DEV Community

Sumitsuke
Sumitsuke

Posted on Edited on

Does AI-generated code silently swallow errors? 120 measured generations: every flagged case was a false positive or a documented fallback

If you gate AI-generated code with linters or a CI rule that hunts for swallowed errors, this experiment suggests the part you actually care about — "is this return None a contract or a cover-up?" — is exactly the part the rule cannot decide.

I started this project convinced that small LLMs routinely swallow failures — catch an error, return an empty value, pretend nothing happened — and that I'd measure the contamination rate. The hypothesis fell apart in a more interesting way than any clean number would have been. This article is the record of a prior being refuted by measurement, with the collapse documented step by step.

TL;DR

Setup: Qwen2.5-Coder 1.5B (Apache-2.0) via Ollama, CPU-only, Windows 11. 12 frozen tasks × 10 generations (seeds 0–9, temp 0.7) = **120 samples: 100 failure-path functions (50 Python, 50 TypeScript) + 20 pure-computation controls. A 7B robustness footnote adds 30 more (150 generations total, all frozen unfiltered in the repo). Classification: Python by AST, TypeScript by regex; **all 120 labels eyeballed, every swallow candidate and boundary case hand-adjudicated* against docstrings, comments, and the function's contract, published as gt.csv.*

  1. A naive Semgrep detector — the kind you might reasonably drop into CI — flagged 4 candidates. Human adjudication: 0 true positives. Correction (2026-09-16): that is only the precision half. The corpus contains 4 samples adjudicated problematic_fallback (py_fetch_json_s2/s3/s7/s9), all of them guard_default, and the shipped rule set reaches none of them — recall 0/4 next to precision 0/4. Correction (2026-09-17): those two fours are different populations. Precision 0/4 mixes languages (2 TypeScript false positives + 2 Python documented fallbacks); recall 0/4 is Python-only. Per language: Python 2 flagged / 0 TP / recall 0/4; TypeScript 2 flagged / 0 TP / recall 0/0 — the TypeScript rows contain no adjudicated problem, so TS recall is undefined, not zero. Read the pair together: "4 flagged, 0 true positives" alone can be read as nothing was there to find, and the corpus says otherwise. Two were false positives (an empty catch in a usage example the model appended outside the function under test), two were legitimate fallbacks, documented (docstring or comment) as the function's contract. But this 0 is not "AI doesn't swallow errors" — my detector targeted try/except and structurally never looked at if-guard default returns, which is where the suspicious shapes actually lived (see finding 3).
  2. In this sample, "fails loudly" and "no handling at all" dominated — not swallowing. Of Python's 50 failure-path generations, 39 (78%) used no try/except: 14 raise explicitly via if/else (loud), 4 return None behind an if guard, 21 just let exceptions propagate. TypeScript went the other way: 33/50 (66%) wrote try/catch with log and/or re-throw — proper handling. The only belief that collapsed was my prior that swallowing would be dominant.
  3. The core finding: whether code "swallows errors" is not decidable from syntax. The same return None shows up as a documented, role-appropriate contract (parse_int: None if it cannot be converted) and as a hazard (fetch_json: 404, 500 — every non-200 collapses into the same None, while a network error raises instead, a third behavior the caller has to know about — with a comment saying so). A comment doesn't make it sound; a bare guard doesn't make it a bug. The deciding information — the function's role, the caller's expectations, the spec — lives outside the pattern. Static rules can surface candidates; they cannot deliver the verdict.
  4. Honest caveats up front: one small model family, 12 tasks, N=10 per task (pseudo-replication — the independent unit is the task, not the generation), single-rater adjudication by me, and a detector whose scope hole I only noticed because the data rubbed my nose in it. Scope: free, local, small models, this corpus. No causal claims about "AI code vs human code" — that comparison belongs to prior work (AIRA, below), not to this experiment.

⚠ Reproducibility scope: generation is nondeterministic and ran once; all 150 generations are frozen unfiltered in the repo. What you reproduce is the deterministic analysis layer — same 120 files in, same distributions and candidate counts out. The final legitimate-vs-swallow labels are human adjudication, published transparently in gt.csv — and the fact that you can't regenerate that layer mechanically is itself the thesis of this article.

Why measure this

AI-generated code quality has a well-cited data point now: AIRA (arXiv:2604.17587, preprint, 2026) ran a deterministic, parser-backed static analyzer over 955 AI-authored and 955 human-authored code samples and reported 1.80× more high-severity findings in AI-authored code (0.435 vs 0.242 per sample), with Broad Exception Suppression (C03) — the "swallowed error" family — as the most frequent check (263 vs 185 in Study 3).

That number gets quoted as "AI code swallows errors." But read the paper closely and it says something more careful, twice:

  • Detection is deterministic and works — the paper even warns that an LLM evaluator misses these suppressions at a rate of 44:1 compared to the deterministic scanner. Pattern-matching is the machine's strong suit.
  • "Flagged ≠ defect." AIRA's own text states that some fail-soft patterns are contextually intentional and that human review is required before remediation. Two of its checks are permanently human-review-only.

So the open question isn't "can we detect suppression patterns?" (yes, deterministically). It's the step everyone skips: can the verdict — legitimate fallback or bug-hiding swallow — be automated too? I couldn't find a published measurement of that gap (pointers welcome). This article measures it with 120 locally generated functions, a naive detector, and every candidate opened by hand.

Scope declaration up front: this is neither "AI is dangerous" nor "static analysis is useless." It's a measurement of where the machine's jurisdiction ends.

Terms, before the numbers

Three things that look identical in grep output and must not be conflated:

  • Default return — on failure, return None / 0 / [] / "" instead of raising. A syntactic shape. Detectable.
  • Silent swallow — a default return that erases information the caller needed: failure becomes indistinguishable from "empty but fine." A semantic judgment.
  • Legitimate fallback — a default return that is the function's contract ("returns None if unparseable"), ideally documented. Also a semantic judgment.

The whole article is about the gap between the first item and the other two. Classification-wise this sits in CWE-703 (Improper Check or Handling of Exceptional Conditions), with CWE-1069 (empty exception block) and CWE-390 nearby; the exception-antipattern literature (de Pádua & Shang, arXiv:1704.00778, Java/C#) maps the same territory.

Why standard linters don't cover this (and what I built instead)

Checked against current docs, not from memory:

  • Python / ruff: S110 (try-except-pass) and S112 exist but are disabled by default — you must opt in via select/extend-select. And a typed swallow like except Exception: return [] sails through anyway.
  • TypeScript / ESLint: no-empty only fires on empty blocks, so catch (e) { return null; } is invisible — it's not empty. no-empty even allows a catch containing only a comment. (Note the cultural assumption baked in there: "has a comment" is treated as "is intentional." Whether the comment describes sound error handling is nobody's department. Hold that thought for the counterexample below.)

So I wrote a Semgrep rule that goes after the meaning-shaped pattern — the naive "before" version, verbatim from the repo:

rules:
  - id: py-swallow-return-default
    languages: [python]
    severity: ERROR
    message: "Exception caught and a default value returned without logging/re-raise (silent fallback)."
    patterns:
      - pattern: |
          try:
            ...
          except $E:
            return $R
      - metavariable-pattern:
          metavariable: $R
          patterns:
            - pattern-either:
                - pattern: None
                - pattern: "[]"
                - pattern: "{}"
                - pattern: "0"
                - pattern: "False"
                - pattern: '""'

  - id: py-except-pass
    languages: [python]
    severity: ERROR
    message: "Exception swallowed with pass (silent failure)."
    patterns:
      - pattern: |
          try:
            ...
          except $E:
            pass
Enter fullscreen mode Exit fullscreen mode

(The TypeScript rules are the same shape — catch { return null }, empty catch — and ship in the repo.)

My hypothesis at this point: "run this over N generations, get a contamination rate." Here is where the measurement starts disagreeing with me.

Setup

  • Tasks: 12, frozen in tasks/tasks.json before generation. Per language: 5 failure-path I/O tasks (config loading, HTTP fetch, env vars, numeric parsing, file head) + 1 pure-computation control (mean/sum — no try needed). Prompts are neutral: no "handle errors" nudging.
  • Task-design disclosure (I did not farm for swallows): pilot runs showed that on plain tasks ("read JSON, return it") the model writes no try/except and nothing interesting happens — which means any "swallow rate" is heavily task-dependent and can be manufactured. I chose failure-inevitable I/O not to inflate the count (the result is ~0 anyway) but to give failure a stage — and kept the pure-computation controls to verify that off-stage, nothing appears. They did: all 20 controls showed zero try/except and zero candidates.
  • Generation: Ollama (MIT), CPU-only, qwen2.5-coder:1.5b (Apache-2.0), temperature 0.7, top_k=40 / top_p=0.9, seeds 0–9 per task, threads fixed at 4. The 7B model appears only in the robustness footnote (N=3 per task).
  • Determinism probe (before trusting anything): at temperature 0 with a fixed thread count, the same seed generated twice was byte-identical (SHA match). Distributions require temp > 0; hence 0.7.
  • Classifier asymmetry, disclosed: Python is classified via ast (a real parse); TypeScript via regex (not a full parser). Cross-language absolute comparisons carry instrument bias — which is why every label was eyeballed and every candidate hand-adjudicated rather than trusting classifier output.
Component Choice
Inference Ollama, CPU (Windows 11)
Model (main) Qwen2.5-Coder 1.5B, default GGUF quantization (Q4_K_M-class)
Model (footnote) Qwen2.5-Coder 7B
Detector Semgrep 1.168.0 (CE) + custom rules above
Analysis Python 3.12, ruff 0.15.12 for the lint baseline

Measured: 2026-07 (corpus generated 2026-07-01). Full per-run metadata in the repo's PROVENANCE file.

Result 1: mostly not swallowing — passing through, or failing loudly

How the 50 failure-path generations per language handled failure:

Failure-path tasks (n=50 per language) Python TypeScript
try/except + log or re-raise (proper) 9 33 (66%)
try/except returning a default (swallow candidate) 2 0
no try/except at all 39 (78%) 17 (34%)

Reading "no try/except = no handling" would be wrong. AST-splitting Python's 39:

  • 14 raise via if/else (e.g. missing env var → raise ValueError(...)) — that's failing loudly, the opposite of swallowing.
  • 4 return None behind an if guard — all four in fetch_json (non-200 → return None). A default return! Which my try/except-scoped detector structurally never saw. Remember these four; they're the article's best specimen.
  • 21 bare — no failure path written; exceptions propagate to the caller.
  • (Check: 14 + 4 + 21 = 39.)

TypeScript's 66% try/catch majority mostly did console.error(...) + throw — textbook handling, no swallowing.

The language difference (Python avoids try/except, TS writes it) is an observation, not a finding: prompt phrasing, language idiom (async/await + try/catch is TS boilerplate), and the AST-vs-regex classifier asymmetry all confound it. The spine of this article is the counterexample below, not this table.

So the original hypothesis — "AI swallows failures at some rate N% I can report" — collapsed in the first table: in this sample, swallowing wasn't the dominant behavior at all.

Result 2: every detector hit was a false positive or documented-legitimate

The naive Semgrep rules flagged 4 of 120 (2 Python, 2 TypeScript). Opening all four by hand:

TypeScript, 2 hits (ts_load_config) = false positives. The function under test was exemplary (code condensed and annotated from the corpus):

async function loadConfig(path: string): Promise<any> {
    try {
        const data = await fs.promises.readFile(path, 'utf8');
        return JSON.parse(data);
    } catch (error) {
        console.error(`Error reading or parsing the file at ${path}:`, error);
        throw error; // logged and re-thrown -- not swallowed (proper)
    }
}
// ...but the model appended a usage example after the function:
(async () => {
    try { const config = await loadConfig('./config.json'); }
    catch (error) { /* Handle any errors... <- comment-only catch */ }
})();
Enter fullscreen mode Exit fullscreen mode

The rule fired on the demo block's catch — (a) outside the function under test, (b) "empty" only because Semgrep's AST ignores comments. Context makes it an obvious false positive; the pattern alone can't know that.

Python, 2 hits (py_parse_int) = documented, legitimate fallbacks. The docstring states the contract — None if it cannot be converted (the second one says the same in a comment).

Bottom line: undisclosed swallowing inside try/except = 0 — and detected problems in the corpus = 0 out of 4 (Python; the TypeScript rows contain no adjudicated problem, so TS recall is 0/0). The first number is precision; the second is recall, and the article originally printed only the first. And now the honest part: my detector's own scope hole. The rules target try/except — the if-guard default returns (the four fetch_jsons) were never in scope. So the corpus does contain default returns; the accurate claim is "zero undisclosed swallows within the detector's scope," not "zero problematic fallbacks in the corpus." Whether those four are problems is precisely the question syntax can't answer — next section.

Result 3 (the core): the same return None, and what separates them isn't in the code

Two functions from the corpus (condensed and annotated). Which one swallows errors?

# (A) numeric parsing: None if unconvertible (contract stated in the docstring)
def parse_int_field(data, key):
    """Returns the int, or None if it cannot be converted."""
    try:
        return int(data[key])
    except ValueError:
        return None            # role-appropriate fallback

# (B) HTTP fetch: None on non-200 (and yes, there's a comment saying so)
def fetch_json(url):
    response = requests.get(url)
    if response.status_code == 200:
        return response.json()
    else:
        return None            # 404 or 500 -- every non-200 flattened into None
                               # (a network error raises instead: a third behavior)
Enter fullscreen mode Exit fullscreen mode

Syntactically, the interesting part is near-identical: on failure, return None. Semantically they're opposites. (A) matches the function's role — "tell me whether this converts" — so None is the answer. (B) permanently destroys the caller's ability to distinguish "no data" from "the fetch failed." And (B) has a comment. Documentation doesn't settle it: a documented return 0 or return None can still silently poison every computation downstream. ("Has a comment = intentional" is exactly the assumption ESLint's no-empty institutionalizes.)

(B) is the strongest specimen this experiment produced: a documented-but-hazardous default return that appeared as a 4-sample cluster (not a one-off), sitting squarely in the blind spot of a try/except-scoped detector — while being exactly the "quietly fails" shape the AIRA numbers gesture at, at population level.

For contrast, non-swallowing code carries its intent inside the syntax:

def get_api_token():
    if 'API_TOKEN' in os.environ:
        return os.getenv('API_TOKEN')
    raise ValueError("The API_TOKEN environment variable is not set.")  # fails loudly
Enter fullscreen mode Exit fullscreen mode

The point, stated carefully: syntactic patterns can surface candidates. The information that separates contract from cover-up — the function's role, the caller's expectations, the spec, and whether the documentation is right — lives outside the pattern. I'm not claiming "undecidable in principle": smarter analysis (types, dataflow, call-site analysis) absolutely narrows the candidates. But "what should this function return, in this context?" is a spec question, and the spec comes from outside the analyzer. Sharper tools shrink the pile; the final reconciliation against intent remains.

Honest scope note: this rests on a small corpus and a handful of specimens — read it as a demonstrated boundary, not a general law about all static analysis. What it demonstrates survives the small N, though, because it's an existence proof: two same-shaped snippets with opposite verdicts, and the verdict-relevant information demonstrably outside the syntax.

What I actually do now: machines generate the candidate list (patterns, distributions, CI notifications — semgrep scan --error exits 1 on candidates, which is fine); a human adjudicates candidates against role, caller, and spec. A green scan is read as "zero candidates for a human to look at," never as "pass."

Where this lands relative to AIRA

Neatly inside its own fine print, it turns out. AIRA detects suppression patterns deterministically and reports the 1.80× population-level difference — candidate generation is the machine's win, and the paper warns that semantic (LLM) evaluation underperforms there, 44:1. But the same paper marks the boundary: flagged patterns aren't necessarily defects, some fail-softs are intentional, human review precedes remediation, and two checks are permanently human-only.

That division of labor is what this experiment probed at specimen level:

  • Machines: sensitivity. Surface default-return shapes, chart try/except usage, gate CI with "candidates exist."
  • Humans: adjudication. Contract or cover-up — read the role, the caller, the spec.

One misquote I need to preempt, because I nearly published it myself (see pitfall 5): AIRA is not a human-evaluation study, and I am not claiming AI code swallows more than human code — that comparison is AIRA's, made with its own methodology, at population level, in a single-author preprint. This experiment neither confirms nor contradicts it; it maps the adjudication residue the paper explicitly leaves to humans.

Everything I got wrong

In the order I got it wrong:

  1. "AI swallows failures; my rule will measure the rate." → Measured: in this sample, swallowing wasn't dominant — pass-through, loud failure, and (in TS) proper handling were. The contamination-rate framing died on the first table.
  2. "The detector's 4 hits = 4 swallows." → All four were false positives or documented-legitimate. The honest number is 0, so I report 0 — with its scope. Don't round adjudication up to detection.
  3. "The false positives just mean my rule is crude; refinement will fix it." → Refinement did kill the TS false positives. It could not touch the real problem: whether a return None is legitimate isn't decided by any pattern, or even by the presence of documentation. This is where the thesis flipped from "measure the rate" to "the verdict doesn't live in syntax." And auditing my own instrument surfaced the scope hole (if-guard defaults) it had from the start. When your tool's errors persist under refinement, suspect the question, not the regex.
  4. The Windows encoding landmine. Semgrep read my TS rule file under Windows' legacy default codepage (cp932) and crashed with UnicodeDecodeError (exit 2) — on an em dash I'd left in a rule message. Fix: ASCII-only rule files + PYTHONUTF8=1 for every run. If your pipeline runs on Windows, pin UTF-8 explicitly or non-ASCII bytes in innocent places will cost you an afternoon.
  5. Writing this article, I did the thing the article warns about. My research notes — AI-assisted — said AIRA judged suppressions "by human evaluation." Pre-publication, I went back to the primary source: it's the opposite. AIRA is a deterministic scanner, and it warns against relying on semantic evaluation for this class. I had a plausible summary, I trusted it, and a "verification discipline" article nearly shipped misstating its key citation's methodology. Corrected against the paper itself. Candidates can come from machines — or from your own notes. The final check against the primary source is the human step, and I almost skipped it.

Limitations

  • One model family, small models (1.5B main / 7B footnote), 12 tasks, one prompt style. Behavior is strongly model- and task-dependent; no generalization claimed beyond this corpus.
  • Statistical independence: N=10 generations per prompt are correlated pseudo-replicates. The independent unit is the task (5 per language), not the generation; don't read the n=50 tables as population estimates. The thesis rides on the counterexample, not the ratios.
  • Detector scope: rules and classifiers target try/except; if-guard defaults are out of scope, so "0 swallows" means "0 undisclosed swallows in scope" — the corpus contains default returns whose adjudication is the whole point.
  • Promise handlers (added 2026-09-20): promise .catch handlers are outside every pattern in the TypeScript rule file, braces or not — every pattern is anchored on a try { ... } catch clause. In this corpus that is 3 of 48 catch-shaped sites (45 try { ... } catch clauses plus 3 promise .catch handlers): ts_fetch_json_s0 and ts_load_config_s7 block-form, ts_fetch_json_s8 expression-form. All three console.error with no return, so 0 of 3 swallow silently; in each file the function's own catch logs and rethrows while the handler logs and stops, so these three are also the corpus's report-without-propagate sites. The numbers above do not move. promise.catch(() => null) — the form the corpus does not contain — is what the rule would miss. Raised by a reader, 2026-09-18.
  • Artefact gap (added 2026-09-16): in results/gt.csv the shape columns subtype and returns_default are filled for the Python rows only — 0 of 60 TypeScript rows carry them (their human_verdict is filled: 31 proper / 27 not_adjudicated / 2 false_positive). Anyone recomputing the cross-tab per language hits 27 unlabelled rows. Raised by a reader who re-derived the analysis from the repo. Partly closed 2026-09-17: for Python rows that have a handler, subtype is no longer blank — it is now guard_default / handler_default / both by where the default return sits (commit 82e5130), so the guard set (subtype in {guard_default, both}, 17 rows) is computable from gt.csv alone. Later the same day (commit f839873) the TypeScript rows got the same three shape columns via tree-sitter, with handling left as the regex verdict the tables above were computed with (the AST agrees on 58/60; the other 2 are the Semgrep false positives). What that exposes rather than closes: the TypeScript guard set is 11 rows (ts_get_item_s1-s9, ts_parse_int_s0/s5) and none of them has been human-adjudicated — their proper comes from the catch logging or throwing, not from anyone reading the default return outside it. TS recall stays 0/0 for that reason, not because the shape is absent — and, more to the point (reader, 2026-09-17), because the TypeScript half has no positive class by construction: the two halves share only 2 of 6 task families, and on fetch_json, the one family that carries all four Python problems, TypeScript is 10/10 proper. Read the TS half as a control arm, not a recall arm. Filling has_raise also moved the four throw-only TS samples from not_adjudicated to loud_fail (same machine rule as Python), so the TS split above now reads 31 proper / 23 not_adjudicated / 4 loud_fail / 2 false_positive.
  • Classifier asymmetry: Python AST vs TypeScript regex (nested-brace catches can slip). Human-verified labels backstop it, but discount cross-language comparisons.
  • 7B robustness footnote (N=3/task, 30 generations): Python 10/15 no-try, 3 proper, 2 candidates; TS 8/15 no-try, 7 proper, 0 candidates. Both candidates: parse_int again, documented — one returns 0 on failure, which is documented and still capable of silently corrupting downstream arithmetic, i.e., the "(B) problem" one more time. N=3, so no robustness claim — only "not refuted in this range."
  • The adjudicator's hole, disclosed next to the detector's: the legitimate-vs-swallow verdicts are single-rater (me), no inter-rater agreement measured. Rubric: (1) does failure yield a default? (2) is that behavior disclosed as intent? (3) does it erase failure information the caller needs? Boundary cases ship with reasons in gt.csv. A second rater is future work — "the last step is human" cuts both ways, so I'm flagging my own last step.
  • Conflict of interest: I do AI-code auditing professionally. "The final call needs a human" is a conclusion that favors my line of work. Mitigation is transparency: frozen tasks, unfiltered generations, published adjudications, and the detector's failures disclosed by me, above.

Reproduce it

git clone https://github.com/sumitsuke/ai-silent-defect-scanner && cd ai-silent-defect-scanner
# Deterministic layer: same 120 files -> same distributions & candidate counts
PYTHONUTF8=1 python scripts/classify_split.py      # failure-path distributions (the n=50 tables)
PYTHONUTF8=1 python scripts/scan_and_count.py      # naive Semgrep candidates (the 4 hits)
PYTHONUTF8=1 python scripts/build_gt.py            # regenerate the adjudication layer, results/gt.csv
make scan-mine DIR=/path/to/your/repo              # candidates for YOUR repo (verdicts are on you)
Enter fullscreen mode Exit fullscreen mode
  • Mechanical reproduction covers distributions and candidate counts. The final labels are my adjudications, published in results/gt.csv with reasons — by design not machine-rederivable (that's the thesis).
  • raw/ ships all 150 generations unfiltered; PROVENANCE records model, quantization, seeds, temperature, thread count, OS, dates. Regenerating gives you a different distribution — that's LLM sampling for you.
  • CPU-only, no GPU needed, no paid APIs anywhere in the pipeline.

Takeaways

  1. Pattern rules find shapes, and the corpus's actual risk shape (if-guard default returns) wasn't the shape my rule watched. Budget for your detector's scope hole, not just its false-positive rate. Correction (2026-09-16): this takeaway originally covered two different boundaries, and only one of them is a defect. (a) The try/except scope is a hole — it costs recall: 100% of the corpus's adjudicated problems sit outside it. (b) The one-statement anchor (handler body is exactly return <default>) is a filter that buys precision — loosening it adds five samples (py_fetch_json_s6, py_first_line_s2, py_load_json_s4, py_load_json_s8, py_parse_int_s2), and all five are has_log=Y and adjudicated proper, 5/5. Keep (b), close (a).
  2. Documentation is not adjudication. A commented return None and a return 0 documented in its docstring both appeared here as disclosed and still hazardous. "Has a comment" is a lint convention, not a verdict.
  3. Split the jurisdictions and staff them accordingly: machine = candidates and distributions in CI; human = contract-vs-cover-up against role, caller, and spec. Green means "no candidates today," not "correct."

If you remember one sentence, please don't make it "static analysis is useless" — the detector did its detection job fine. Make it: syntax can surface the suspects, but conviction requires reading intent — and intent isn't stored in the AST.


Detection code, all 150 generations, and the human adjudication layer (gt.csv): github.com/sumitsuke/ai-silent-defect-scanner. Every number above is measured; anything not measured is labeled as not measured.

This is an English adaptation of my Japanese article on Qiita (Qiita is a Japanese dev-blogging platform) — written by me in Japanese, restructured and translated with AI assistance, human-reviewed. If you spot an error, comments and issues are open; I'll verify against the frozen corpus and correct with a changelog.

Verification record (environment, verdict, last verified date, evidence) and the canonical write-up: https://sumitsuke.jp/via/devto/lab/ai-code-silent-fallback/ — code, data and reproduction: https://github.com/sumitsuke/ai-silent-defect-scanner.


Corrections

2026-09-16 — After publication, a reader re-derived the deterministic layer from the repo and pointed out that the report card was missing its recall half. I recounted from results/gt.csv and confirmed every checkable claim: 4 adjudicated problematic_fallback samples, all guard_default, none reachable by the shipped rules (recall 0/4); 9 of the 20 controls carry a default return; the five samples added by loosening the one-statement anchor are 5/5 has_log=Y and proper; the TypeScript rows carry no shape labels. One number I could not reproduce from the labels alone is their precision 4/17 — from gt.csv I reach 13 (returns_default=Y AND handling=no_try_except); the other four would have to come from reading the source. Changes above: the recall row in the TL;DR and in Result 2, the split of scope-hole vs anchor in Takeaway 1, and the artefact gap in Limitations. Nothing in the measured data changed.

2026-09-17 — The same reader came back with two things. First, a correction of their own: "precision 0/4, recall 0/4" reads like one 2×2 but is not — the 4 flagged are 2 TypeScript + 2 Python, the 4 problems are all Python, and TypeScript has no adjudicated problem at all (recall 0/0). I had printed the pair the same way; language is now attached in the TL;DR and Result 2. Second, the 4/17 I said I could not reproduce from the labels: they posted a 30-line AST detector (Return of a default value outside any except handler), and I ran it against raw/python/ — it fires on exactly 17 files, the four "extra" ones are py_fetch_json_s6:13, py_parse_int_s3:22, py_parse_int_s6:14, py_parse_int_s9:12, line numbers as they stated, and the verdicts on the 17 are 4 problematic_fallback / 2 legit_fallback / 2 proper / 9 controls (not_adjudicated) — recall 4/4, precision 4/17. The reason the labels could not reach those four is that gt.csv recorded whether a default return exists but not where; their proposed column closes it, and build_gt.py now writes guard_default / handler_default / both for every Python row that has a handler (of the four: py_fetch_json_s6 and py_parse_int_s3/s9 are both, py_parse_int_s6 is guard_default — its handler raises, and its only default return is in the else branch). Nothing in the measured data changed; one column was added — and, by evening, extended to the TypeScript rows (f839873), which turned the TS artefact gap from "no shape columns" into "11 guard-default samples nobody has adjudicated" (Limitations).

2026-09-17, later — The reader re-ran both fixes against the source (Python 17/60 and the four line numbers reproduce; the TypeScript guard set of 11 reproduces exactly) and added two things I had not stated. First, the "4 flagged" row is not one detector's output: the two Python hits come from py-swallow-return-default, the two TypeScript hits from ts-empty-catch (I confirmed this from scan_and_count.py's hit list); the return-default rule fires on 0 TypeScript files, because its one-statement anchor excludes all 9 TypeScript catches that log first — every one of them is console.error(...) then return null/undefined, and every one is proper. Second, the TypeScript half is best read as a control arm: it shares only two tasks with Python, and the task that carries all four Python defects (fetch_json) is 10/10 proper in TypeScript. (I originally wrote here that the TypeScript half "has no positive class by construction of the task sample" — withdrawn on 2026-09-18, see below; whether it has a positive class is an adjudication outcome, not a property of the sample.) The two outstanding rulings on ts_parse_int_s0/s5 remained open at this point. The reader asked to stay unnamed.

2026-09-18 — The same reader's third round. Two things. (1) The sentence "no positive class by construction of the task sample" was too strong and is withdrawn: the guard filter fires on the same two shapes in both halves; TypeScript's guard rows are 9 proper (machine label: the catch logs or throws) and 2 that were open. (2) I ruled the two open rows, ts_parse_int_s0/s5: no catch in either file, and the docstrings state the default ("it returns null" / "otherwise undefined") — the same rule that made py_parse_int_s3/s9 legit_fallback, so both are legit_fallback (commit in the repo, review_note filled). The TypeScript recall line therefore still has no positive to count, but the honest reading is now "0 of 2 opened, both documented", not "0/0, nothing to open". The Python half is not fully adjudicated either: with subtype ∈ {guard_default, both} and returns_default = Y it is 17 rows — 4 problematic / 2 proper / 2 legit / 9 not_adjudicated. One more correction of my own number in the comments: "49 of 60 TypeScript files end with a demo block" was the wrong reading (49 = files with a console.* call outside any catch); files whose last top-level statement is a console.* call are 22/60.

2026-09-20 — The same reader's fourth round: a scope hole the brace-shaped census could not see. Promise .catch handlers match nothing in the TypeScript rule file, with or without braces; in this corpus that is 3 of 48 catch-shaped sites (45 try/catch clauses plus 3 promise handlers), and all three are console.error with no return (0 of 3 swallow silently; all three log and stop where the function's own catch rethrows), so no number above changes. Stated in Limitations now, next to the if-guard hole, with promise.catch(() => null) as the one-line form the rule would miss.

2026-09-21 — Same reader, fifth round: the 48 had no instrument behind it and reads as 45 to anyone recomputing from the rule's anchor. It is now stated as 45 try/catch clauses plus 3 promise handlers, in Limitations and here; and the three handlers are named as what they are in this corpus — the report-without-propagate sites, where the function's own catch rethrows and the call-site handler logs and stops. No number moves.

Top comments (48)

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

I re-derived your deterministic layer against the frozen corpus before writing anything. One caveat on method: semgrep isn't installed on this machine, so I re-implemented the two Python rule shapes as AST matches rather than running your tool — with the same one-statement anchor (handler body is exactly return <default>) it reproduces exactly your two Python hits, py_parse_int_s3 and py_parse_int_s9. Everything below is computed from raw/ plus results/gt.csv at 18b9a19, so you can disagree with any cell.

The detector's report card is missing its second half. You report the flagged side: 4 candidates, 0 true positives. Your own labels say the corpus contains 4 samples adjudicated problematic_fallback — py_fetch_json_s2/s3/s7/s9 — and the shipped rule set flags none of them, because all four are guard_default, out of scope. So the row is precision 0/4 and recall 0/4. "Zero undisclosed swallows in scope" is exactly right as stated; the number that makes it readable next to the other one is "four adjudicated problems in the corpus, zero detected". As published, "4 flagged, 0 true positives" can be read as nothing was there to find.

The scope hole has a size, and it costs one rule. Same corpus, same tool class: flag any default return that is not inside an except. Over the 60 Python files that fires on 17 — the 4 problems (recall 4/4), the 9 controls, 2 proper, 2 legit_fallback — precision 4/17 = 24%, i.e. one true positive per 4.25 candidates a human has to read. So 100% of the corpus's problems sit in the shape you name as out of scope, and a rule of the same kind surfaces all of them. That is your takeaway #1 with a number attached, and it makes the repair look cheap: the cost of closing the hole is 13 extra candidates to adjudicate, not a different method.

The controls are the densest source of that same shape. 9 of your 20 controls carry a default return by your own labelling (if not numbers: return 0 # Return 0 for an empty list to avoid division by zero), and the guard-scoped rule fires on 9/9 of them while the try/except rule fires on 0/20. So "all 20 controls showed zero try/except and zero candidates" is a scope-bound statement: off-stage is not empty, it holds 9 of your 13 guard_default rows. They are also return 0 rather than return None, which your 7B footnote already names as the (B) problem one more time — documented, and still able to corrupt downstream arithmetic. The arm built to be the negative case carries the shape that can't be verdicted from syntax, which is an argument for your thesis rather than against it.

Your anchor is a feature, and your labels say so — but it is a second boundary, not the same one. The rule requires the handler body to be exactly return <default>. Loosen it to "the handler contains a default return anywhere" and it fires on 5 more samples: py_fetch_json_s6, py_first_line_s2, py_load_json_s4, py_load_json_s8, py_parse_int_s2 — all five have has_log=Y and all five are adjudicated proper. So the strict anchor is, on your own data, a "no log ⇒ silent" test with 5/5 agreement. There are two boundaries around this detector and they are different in kind: the try/except scope is a hole, the one-statement anchor is a correct filter. "Budget for your detector's scope hole" currently covers both, and the numbers say one of them you'd want to keep.

One artefact note. gt.csv's per-sample shape columns are effectively Python-only: subtype and returns_default are filled for 0 of 60 TypeScript rows (the verdicts are filled — 31 proper, 27 not_adjudicated, 2 false_positive). Your reproducibility note discloses the AST-vs-regex asymmetry; this is where it lands in the artefact — a reader computing the same matrix per language, or looking for the same shape in TS, arrives at 27 unlabelled rows.

Happy to hand over the ~30-line detector and the exact cross-tab if you want to fold the recall half into the repo.

Collapse
 
tauridev profile image
Sumitsuke •

You went and re-derived it, so the least I can do is check every cell against the same file before answering. I did, at results/gt.csv in the repo, and the parts that live in the labels all come out the way you say:

  • problematic_fallback = 4, and they are exactly py_fetch_json_s2/s3/s7/s9. All four have subtype=guard_default, so the shipped try/except rule cannot reach them. Recall 0/4 alongside precision 0/4 — you're right, and that row is missing from the article.
  • 9 of the 20 controls carry a default return (returns_default=Y), and 9 of the 13 Python guard_default rows are controls. The negative arm is the densest source of the shape the detector can't verdict. That's a better sentence than anything in my piece.
  • The strict one-statement anchor is a filter, not a hole. Your five — py_fetch_json_s6, py_first_line_s2, py_load_json_s4, py_load_json_s8, py_parse_int_s2 — are has_log=Y and adjudicated proper, 5/5. I had them collapsed under one "budget for the scope hole" takeaway; they're two different boundaries and only one of them is a defect.
  • The TS rows are unlabelled for shape: subtype and returns_default are empty for 60/60 TypeScript rows, while human_verdict is filled (31 proper / 27 not_adjudicated / 2 false_positive). Anyone recomputing the cross-tab per language hits that wall, and the artefact should say so.

One number I could not reproduce from the labels alone: your 17. From gt.csv I get to 13 by taking returns_default=Y AND handling=no_try_except (the 4 problems + 9 controls). The remaining 4 — 2 proper, 2 legit_fallback — must come from reading the source rather than the labels, since those rows are marked as handled inside an except. So I can confirm recall 4/4 and the shape of the trade, but not precision 4/17 from the artefact as published. If your detector fires on those four, that's a fact my own labels don't carry, which is itself worth recording.

What I'm changing, and I'd rather say it here than quietly:

  1. The report card gets its second half — 4 adjudicated problems in the corpus, 0 detected printed next to "4 flagged, 0 true positives". As published, mine reads as "nothing was there to find," which is not what the corpus says.
  2. The scope hole and the one-statement anchor get separated. One costs recall, the other buys precision; they were in the same sentence.
  3. The TS labelling gap goes in the artefact note, not just the reproducibility note.

Yes to the ~30-line detector and the cross-tab. If you send it, I'll run it against the same corpus, publish the confusion matrix in full — both halves this time — and credit the recall half to you in the repo and the article.

For what it's worth, this is the second time today the same failure has bitten me from a different direction: a check that only ever sees the cases you hand it, staying quiet, and the quiet being read as an answer. Yours is the sharper instance, because the cases it never sees were sitting in my own labels the whole time.

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

You checked the cells, so let me check mine back — starting with one of my own, because it does not survive the same treatment.

My "0/4 and 0/4" mixed two different fours. The article's 4 flagged = 2 TypeScript (ts_load_config_s0/s2, both false_positive) + 2 Python (py_parse_int_s3/s9, both legit_fallback). My recall denominator is the 4 Python problematic_fallback rows. So one side of that sentence is language-mixed and the other is Python-only. Per language the shipped rule is: Python 2 flagged / 0 TP; TypeScript 2 flagged / 0 TP — and TypeScript has zero adjudicated problems, so its recall is 0/0, undefined rather than 0. Written as "precision 0/4, recall 0/4" it looks like one clean 2×2, and it is not one. Precision 0/4 is right; the recall half needs its language attached.

Your question — the 17. You are right that 13 come out of the labels and 4 do not. The 4 are:

sample handling subtype verdict guard rule fires at
py_fetch_json_s6 proper (blank) proper line 13
py_parse_int_s3 swallow_cand (blank) legit_fallback line 22
py_parse_int_s6 proper (blank) proper line 14
py_parse_int_s9 swallow_cand (blank) legit_fallback line 12

Why the labels cannot reach them: handling describes what the handler does, and returns_default says a default return exists somewhere. There are 8 rows that have both a handler and returns_default=Y, and subtype is blank on all 8 — the artefact does not record where that default return sits. My rule fires on 4 of those 8 (the table) and not on the other 4 (py_first_line_s2, py_load_json_s4, py_load_json_s8, py_parse_int_s2), whose only default return is the handler's. So 17 = 13 + 4, and the 4 are exactly "a second, non-handler default return coexisting with a handled try/except" — a shape your column set currently has nowhere to put.

One column closes it. If subtype stopped being blank whenever a handler exists — say guard_default / handler_default / both — then the guard set is subtype in {guard_default, both} and both matrices become computable from gt.csv with no source reading. py_parse_int_s3/s9 would be both, which also records something the current artefact hides: those two rows are the entire overlap between the two rules. Everywhere else they disagree.

The cross-tab, as requested (Python, my re-implementation, not semgrep):

rule flagged TP FP precision recall
shipped try/except 2 0 2 0/2 0/4
guard (this) 17 4 13 4/17 = 23.5% 4/4
union 17 4 13 4/17 4/4

Per-sample, guard rule only — sample | handling | subtype | verdict | lines:

py_avg_control_s0|no_try_except|guard_default|not_adjudicated|[3]
py_avg_control_s1|no_try_except|guard_default|not_adjudicated|[3]
py_avg_control_s2|no_try_except|guard_default|not_adjudicated|[3]
py_avg_control_s3|no_try_except|guard_default|not_adjudicated|[3]
py_avg_control_s4|no_try_except|guard_default|not_adjudicated|[7]
py_avg_control_s5|no_try_except|guard_default|not_adjudicated|[3]
py_avg_control_s6|no_try_except|guard_default|not_adjudicated|[3]
py_avg_control_s8|no_try_except|guard_default|not_adjudicated|[4]
py_avg_control_s9|no_try_except|guard_default|not_adjudicated|[12]
py_fetch_json_s2|no_try_except|guard_default|problematic_fallback|[14]
py_fetch_json_s3|no_try_except|guard_default|problematic_fallback|[15]
py_fetch_json_s6|proper|(blank)|proper|[13]
py_fetch_json_s7|no_try_except|guard_default|problematic_fallback|[14]
py_fetch_json_s9|no_try_except|guard_default|problematic_fallback|[15]
py_parse_int_s3|swallow_cand|(blank)|legit_fallback|[22]
py_parse_int_s6|proper|(blank)|proper|[14]
py_parse_int_s9|swallow_cand|(blank)|legit_fallback|[12]
Enter fullscreen mode Exit fullscreen mode

The detector (Python 3.10+, no dependencies) — python guard_default.py raw/python/*.py reproduces the 17 lines and the four line numbers in the table above:

import ast, sys

def is_default(node):
    if isinstance(node, ast.Constant):
        return node.value is None or node.value is False or node.value == 0 or node.value == ""
    if isinstance(node, (ast.List, ast.Dict, ast.Tuple)):
        elts = getattr(node, "elts", None) or getattr(node, "keys", [])
        return len(elts) == 0
    return False

class GuardScan(ast.NodeVisitor):
    def __init__(self, path):
        self.path = path

    def visit_ExceptHandler(self, node):
        return  # handler bodies belong to the try/except rule, not this one

    def visit_Return(self, node):
        if node.value is not None and is_default(node.value):
            print("%s:%d guard-return-default" % (self.path, node.lineno))
        self.generic_visit(node)

for path in sys.argv[1:]:
    GuardScan(path).visit(ast.parse(open(path, encoding="utf-8").read()))
Enter fullscreen mode Exit fullscreen mode

Two scope statements about that table. It is Python only — my TypeScript stand-in for the try/except rule is a regex, not semgrep, and it fires on 5 files where yours fires on 2, so I claim nothing at all on the TS side. And that is not only my limitation: with subtype and returns_default empty for 60/60 TS rows, nobody can compute TS recall from the artefact even if the detector were exact. The rule firing on the 9 controls is the point, not a defect — it is a candidate generator, and the negative arm is where the shape is densest (9 of your 13 Python guard_default rows).

Three things I would keep exactly as you plan them, and one I would add: the column, since it is what lets a reader recompute both halves instead of taking either of us on our word. Credit is yours to give; what I want out of it is that the cross-tab is checkable by someone who has only the corpus.

Your last paragraph is the part I would keep. A check that only sees the cases you hand it, staying quiet, and the quiet read as an answer — that is a nicer statement of the same defect than the one I posted.

Thread Thread
 
tauridev profile image
Sumitsuke •

Ran your detector against raw/python/ before writing this: 17 files, and the four line numbers in your table (fetch_json_s6:13, parse_int_s3:22, parse_int_s6:14, parse_int_s9:12) come out exactly as stated. So the 4/17 I said I could not reproduce from the labels is reproduced from the source, and you were right about why — gt.csv recorded whether a default return exists, not where.

The column is in. build_gt.py now writes guard_default / handler_default / both for every Python row that has a handler (commit 9d7bf7f in the repo, README documents the values). Regenerated gt.csv: guard set = subtype in {guard_default, both} = 17 rows, verdicts 4 problematic / 2 legit / 2 proper / 9 controls — your cross-tab, from the artefact alone. One row to note: py_parse_int_s6 comes out guard_default rather than both, because its handler raises and the only default return is in the else branch; fetch_json_s6 and parse_int_s3/s9 are both, as you said. And yes — s3/s9 being the entire overlap between the two rules is now visible in the file instead of only in your comment.

Your own correction I have taken as well. The article now attaches language: Python 2 flagged / 0 TP / recall 0/4; TypeScript 2 flagged / 0 TP / recall 0/0, undefined. The 2026-09-17 entry in Corrections credits the detector and the column to a reader; if you would rather be named, say so and I will put the handle in.

Still open on my side: the TypeScript rows. 60/60 without shape columns means nobody can do for TS what you just did for Python, and my TS classifier is a regex. That is the next thing to fix in the repo, not something I can close in a reply.

Thread Thread
 
howcani_howcani_77e786a89 profile image
howcani howcani •

I ran your fix against the source before answering, and it holds on every point I can check.

At 89bfa35: my detector fires on 17/60 files with the four line numbers I published (13 / 22 / 14 / 12). Your declared set, subtype in {guard_default, both}, is those same 17 rows — the symmetric difference is empty in both directions, and the verdict mix is 4 problematic / 2 legit / 2 proper / 9 controls, identical to the set my detector produces. py_parse_int_s6 as guard_default is right and I would have had it wrong: the handler raises at line 11, so the only default return is the one in the else at 14, and there is no handler default to be "both" with. And the four rows your column calls handler_default are exactly the four my rule does not fire on, which is the part I could previously only assert from reading source.

Then you closed the TS half six minutes after writing that comment (f839873), so let me do the thing you said nobody could do yet — the TS cross-tab. Same disclaimer as last time: semgrep still isn't installed here, so I re-implemented both shapes with tree-sitter.

  • The guard set reproduces exactly: 11 rows, ts_get_item_s1-s9 + ts_parse_int_s0/s5, symmetric difference empty.
  • The two TS hits of the shipped rule also reproduce — but from the other rule in the same file. ts-empty-catch hits ts_load_config_s0 L23 and ts_load_config_s2 L24, the comment-only catch in the demo block, which are precisely the two rows a human ruled false_positive. The return-default rule hits 0 files on the TS half.

Which makes the anchor's effect language-dependent, and that is the thing I did not see last round. The anchor says "the catch body is exactly return <default>", and on both sides it excludes exactly the rows that log first:

strict anchor loose (any default return in the catch) what the anchor excluded
Python 2 (parse_int_s3/s9, legit_fallback) 7 5 rows, all has_log=Y / proper (5/5)
TypeScript 0 9 9 rows, all proper

In all nine excluded TS rows the catch body is exactly two statements — console.error(...) then return null (or undefined) — so on that half the anchor removes 100% of the return-default catches, by the same rule it uses on Python. So the published "4 flagged (2 Python + 2 TypeScript)" is not one detector's output: the Python two come from the return-default rule, the TypeScript two from the empty-catch rule, and the rule that reaches the corpus's problems is neither. That row's 0 is a union, and each half of it is zero for a different reason.

On the TS work you say is next — the column is in, and the TS cross-tab still reads 0/11 with recall 0/0, and I'd argue that is not a labelling gap. Two numbers from the regenerated file:

  • The TS half is more adjudicated than the Python half — 37/60 rows carry a verdict against 29/60 — and still contains zero problematic_fallback.
  • The two halves share only 2 of 6 task families (fetch_json, parse_int). On fetch_json, the single task family carrying all four adjudicated problems on the Python side, TS is 10/10 proper.

So the TS half's positive class is empty by construction of this corpus, and no column, classifier or adjudication round moves recall 0/0. The two honest options I can see: add problem-bearing TS generations, or print the TS half as a control arm rather than a recall arm — it largely is one already (10/10 proper on the shared task, and the only two false_positive rows in the corpus). That is the same shape as the fix you just made, one level up: the Python recall row was missing its denominator; the TS row has neither numerator nor denominator, and the reason lives in the task sample rather than in the schema.

For the record the TS guard set is 9 proper + 2 not_adjudicated (ts_parse_int_s0/s5), so its precision is 0/11 today with two rulings still outstanding — and those two look like the same shape as the other nine, which is what makes the control-arm reading cheap rather than a way of avoiding a verdict.

Please leave it as the reader. Nothing in the cross-tab needs a name to be recomputed, and the entry is more useful if it can be checked without knowing who wrote the sentence. Taking the language split into Corrections is the right home for it — that half was mine and it was wrong.

Thread Thread
 
tauridev profile image
Sumitsuke •

Checked before replying, as before. Both of your new points hold against the repo.

The "4 flagged" row: the two TypeScript hits come from ts-empty-catch (ts_load_config_s0/s2, the comment-only catch in the demo block) and the two Python hits from py-swallow-return-default — scan_and_count.py's hit list says exactly that, and I had been printing the union as if it were one rule. I also read the nine TypeScript catches the strict anchor excludes: every one is console.error(...) followed by return null or undefined, all proper, so on that half the anchor removes 100% of the return-default catches. Both facts are now in Corrections, the README, and the Japanese Lab page.

The control-arm reading I accept, and it is the better description of what the TypeScript half is: 2 of 6 task families shared, and on fetch_json — the only family carrying Python's four problems — TypeScript is 10/10 proper. I have written it that way in Limitations: no positive class by construction of the task sample, so no column or adjudication round can move recall off 0/0. The two outstanding rulings on ts_parse_int_s0/s5 stay open; I agree they look like the other nine, but those are mine to make with the same rubric as the Python rows, not something to settle in a comment thread.

Unnamed it stays — the entry reads "a reader" and the cross-tab checks without knowing who wrote it. Thank you for doing the TS half before I got to it.

Thread Thread
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Reproduced, both directions. e5b626a3 carries the two facts in the README (the union sentence and the control-arm reading), scan_and_count.py's hit list does attribute the four candidates 2 + 2, and my scanner gets the same exclusions you quote: 5 Python and 9 TypeScript bodies excluded by the strict anchor, all 9 of the TS ones console.error(...) + return null/undefined, and ts_fetch_json 10/10 log=Y raise=Y. The nine-for-nine match is why the next part is worth your time.

One sentence in that README is mine and it is wrong. "No positive class by construction of the task sample" — I wrote that last round and you adopted it into Limitations. It is too strong, and the two rows you left open are the reason.

ts_parse_int_s0 and s5 carry no catch at all. Your own gt.csv says handling=no_try_except for both, and a brace scan of the two files finds 0 catch bodies; s0 is if (typeof value === 'string' && !isNaN(Number(value))) return parseInt(value); followed by return null;. The nine ts_get_item_* rows are the other shape — catch { console.error(...); return undefined }, subtype=both, handling=proper — so "they look like the other nine" applies the rubric's premise, that there is a catch to judge, to two rows that have none. They look like py_fetch_json_s2/s3/s7/s9: same filter (no_try_except + returns_default=Y), four rows, all four problematic_fallback.

But that filter has a second half, and it is the half that fits these two:

family default-returning rows, failure-prone half verdict
fetch — fallback erases 404/500/network/empty Py 4 problematic 4/4
fetch — every row logs and throws TS 10 proper 10/10
parse — fallback is what the docstring promises Py 2 legit 2/2
parse TS 2 open

py_parse_int_s3/s9 are legit_fallback with the reason written beside them: "docstring documents 'None if key missing or not convertible' = intended contract", "comment 'Return None if conversion fails' documents intent". s0's docstring says "If the conversion is not possible or the value is not a number, it returns null", s5's says "otherwise undefined". The verdict follows the task, not the shape, the language, or the log — which is the article's own thesis, and these two rows are where it can be shown inside the TS half instead of argued by analogy from the Python half.

What that changes is smaller than "recall moves" and larger than bookkeeping. 0/0 is undefined, and an undefined ratio cannot tell a reader whether the corpus was searched; 0/2 says two rows were opened and both were documented. It also puts the two rules on one footing: in the failure-prone halves the guard filter fires on 12 Python rows and 12 TypeScript rows. Python's twelve are fully adjudicated (4 problematic / 6 proper / 2 legit); TS's are 10 proper and 2 open. The entire distance between "recall 4/4" and "recall 0/0" is those two adjacencies — while the catch-body rule has no TS target by construction, which you now say in Limitations and I agree with: 45 TS catch bodies, 40 carry a log, and the only 2 silent ones are the demo block.

Not a ruling from me — they are yours and you have the two candidate precedents in the same file now, each one filter away.

On the neighbouring thread's 49/60: I can hit 49 exactly, but under a different reading than the sentence. console. appearing outside any catch body → 49/60. "Ends with a demo block that logs the return" → 22/60 (last top-level statement is a console.*( call), 25/60 (the tail after the last column-0 } contains one), 31/60 (any column-0 statement after it), 55/60 (contains console. anywhere). So the number is real but means "has a non-catch console call", which sweeps in check()-style helpers and logs sitting inside guards — roughly two dozen files that are not a trailing demo. Your conclusion there doesn't depend on it; the share of the TS half that is the generator talking to itself does.

Instrument: my own brace/AST scan on the frozen raw/ at e5b626a3, per-language body shapes and the four trap counts; the 12/12 and the family table come from your gt.csv columns, so they are one filter away from being yours. No semgrep here, and I did not touch the two open rows.

Thread Thread
 
tauridev profile image
Sumitsuke •

Checked against the repo before answering, same as before. Your facts hold: ts_parse_int_s0 and s5 carry no catch (gt.csv handling=no_try_except, subtype=guard_default, verdict not_adjudicated for both); s0's docstring says "If the conversion is not possible or the value is not a number, it returns null", s5's says "otherwise undefined"; and py_parse_int_s3/s9 are legit_fallback for exactly that reason — the docstring or comment states the default as the contract. So the sentence you flagged is mine to withdraw. "No positive class by construction" was wrong on the word "construction": whether the TypeScript half has a positive class is an adjudication outcome, not a property of the sample, and two rows were still open when I wrote it. The accurate statement is: the guard filter fires on the same two shapes in both halves; TypeScript's rows were 9 proper and 2 open at that point; the catch-body rule has no positive TypeScript target (its two hits are the demo-block catches, per the earlier round); and "0/0" on the TypeScript line should read "0 of 2 opened", which is a different — smaller — claim than "nothing to find". That wording is now in the Corrections entry (2026-09-18) and the repo README in place of mine.

One row of your table I could not reproduce from gt.csv as it stands, so I will not adopt it: "Python 12, fully adjudicated (4 / 6 / 2)". With subtype in {guard_default, both} and returns_default=Y I get Python 17 rows — 4 problematic / 2 proper / 2 legit / 9 not_adjudicated — and TypeScript 11 (9 proper / 2, ruled below); with handling=no_try_except and returns_default=Y it is Python 13 (4 problematic / 9 not_adjudicated) and TypeScript 2. Either way the Python half is not fully adjudicated either: nine of its guard rows are open, which the two TypeScript ones no longer are. If your 12 comes from a filter I am not seeing, name it and I will run it; the "12 and 12 on one footing" symmetry is nicer than what the file supports today.

On the two rows themselves: I read them the same way you do — same filter, same documented default, same shape as py_parse_int_s3/s9 — and I ruled them, in the file rather than here: both legit_fallback, intent_disclosed=docstring, review_note naming the precedent (commit 77eddf3; gt.csv changed in exactly those two rows). TypeScript's guard rows are now 9 proper / 2 legit / 0 open; the TypeScript recall line still has no positive to count, and reads "0 of 2 opened, both documented".

And one correction I owe the other thread, since you measured it: "49 of 60 end with a demo block" was the wrong reading of my own count. 49 is "has a console.* call outside any catch". "Last top-level statement is a console.* call" is 22/60 (I get 22 with a column-0 scan; 55/60 contain console. at all). The conclusion there — that these are the generator demoing itself, not consumers — does not depend on which number, but the number I wrote was the wrong one, and I've said so in that thread and in the Corrections entry.

Thread Thread
 
howcani_howcani_77e786a89 profile image
howcani howcani •

The filter is returns_default == Y, scoped to the six task families — the control arm is excluded. At 77eddf3 that is Python 12 and TypeScript 12, and it lands on the same number in both halves for the same reason: returns_default is the guard rule's own predicate, not a shape column, and the control never enters it (ts_sum_control is ten rows of bare / returns_default=N).

python      returns_default=Y, task != control   4 problematic / 6 proper / 2 legit
typescript  returns_default=Y                    10 proper / 2 legit      (10 / 2 open at e5b626a3)
Enter fullscreen mode Exit fullscreen mode

The nine not_adjudicated rows both of your filters kept are py_avg_control_s0..s6, s8, s9 — nine rows, all guard_default, all un-adjudicated by design. That is the whole difference between the three numbers:

  • handling=no_try_except & returns_default=Y → 13 = the four py_fetch_json problematic rows + those nine. It cannot see py_fetch_json_s6, py_first_line_s2, py_load_json_s4/s8 or py_parse_int_s2/s3/s9, which all carry a try.
  • subtype in {guard_default,both} & returns_default=Y → 17 = eight of the twelve (it drops the four handler_default rows) + the same nine.

So "the Python half is not fully adjudicated either" is a property of the control arm, not of the census: those rows are open because they are the control, and no filter that keeps them can converge. Filter on the predicate and the two halves read 12 and 12 with no open rows on either side.

Your withdrawal is right and the replacement is the stronger sentence. 0 of 2 opened, both documented differs from 0/0 in exactly the way that matters — an undefined ratio cannot tell a reader whether anything was searched, and the file now supports the defined one: both rows are legit_fallback with the precedent named beside them (py_parse_int_s3/s9, same filter, same documented default).

Your three corrected numbers reproduce here from the frozen raw/ with an independent brace scan, all three exactly: 49/60 (a console.* call outside any catch body), 22/60 (last top-level statement is a console.*(...) call), 55/60 (contains console. anywhere). I'll take the reading — the 49 was "has a non-catch console call", which sweeps in the guard-style logs and the check() helpers.

One number I could not reproduce, so I am not adopting it: the catch-body totals. My scan gets 47 bodies, 43 containing a console. call, 2 silent — against your 45 / 40 / 2. The load-bearing half does reproduce, including the location: the two silent ones are comment-only catches in ts_load_config_s0.ts and s2, and both sit in the demo IIFE, while the function's own catch in that same file is console.error(...) + throw. (Two more of the 47 are throw-only with no log — ts_load_config_s3, s9 — which are re-raises rather than swallows under the article's own two columns.) 47 ≠ 45 on one frozen tree, so the totals stay yours and unconfirmed here; the silent set and its address are what I could verify.

Thread Thread
 
tauridev profile image
Sumitsuke •

Your filter reproduces here at 77eddf3, and I am adopting the reading. returns_default == Y with the control families excluded gives Python 12 = 4 problematic / 6 proper / 2 legit and TypeScript 12 = 10 proper / 2 legit, zero open rows on either side. The nine rows my two filters kept open are py_avg_control_s0-s6, s8, s9 — the control arm, un-adjudicated by design — so "the Python half is not fully adjudicated either" was a statement about the control, not the census, and I withdraw it. Your decomposition of my 13 and 17 is exact (the four py_fetch_json rows plus those nine; eight of the twelve plus those nine, the four dropped being the handler_default rows).

On the catch bodies, the two-body gap is the promise form. A brace scan of the frozen raw/ gives 45 try/catch bodies: 41 with a console. call, 2 comment-only (ts_load_config_s0 and s2, both in the demo block at the bottom of the file — an async IIFE in s0, a top-level try in s2 — while the function's own catch in those files logs and rethrows), 2 throw-only (ts_load_config_s3 and s9). There are also three promise .catch handlers: two block-form, ts_fetch_json_s0 and ts_load_config_s7, both console.error only, and one expression-form one-liner in ts_fetch_json_s8. Counting the two block-form handlers as bodies gives your 47 / 43 / 2, so the two scans agree on one tree once that choice is stated. On 45 / 40 / 2: I could not find that triple in the article, the README, or my replies, so I cannot defend the 40 — my scan says 41 — and the numbers I stand behind are 45 / 41 / 2 for try/catch bodies, plus the two block-form promise handlers.

Thread Thread
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Your reconciliation is exact, and the cause is my instrument rather than a difference of reading. I split the two forms on the frozen raw/:

try { occurrences (bodies)         45
catch-clause occurrences           48   <- includes the 3 promise handlers
promise .catch(...) total           3
  block-form (braces)               2   ts_fetch_json_s0, ts_load_config_s7
  expression-form (no braces)       1   ts_fetch_json_s8
Enter fullscreen mode Exit fullscreen mode

My pattern was \bcatch\b[^{]*\{, and \b matches after the dot in .catch(, so it counted the method name as if it were the clause: 45 bodies + the 2 block-form handlers = my 47, and 41 logging bodies + the same 2 = my 43. Silence 2 = 2 in both. One tree, one scan each, once that choice is stated — and the choice was made by a regex, not by me.

On the triple: you are right that it is not findable, and it was mine to attribute correctly. It came from your own earlier comment on this thread ("45 TS catch bodies, 40 carry a log, and the only 2 silent ones are the demo block"), and I compared my scan against that sentence without checking it against the tree. Your corrected 45 / 41 / 2 is what the files support and I have no number of my own to add to it.

The part I would keep from this is the form neither of us counted. ts_fetch_json_s8's handler is .catch(error => console.error('Failed to fetch data:', error)) — no braces, so it is not a body in your 45 or in my 47, and every pattern in rules/silent-fallbacks-typescript.yml is anchored on a try/catch clause (try { ... } catch ($E) { return null; } and nine siblings, plus the empty-catch pair). So a promise handler matches nothing in that rule file whether it has braces or not, and promise.catch(() => null) — a fallback that erases the failure, which is the subject of the article — has no catch ($E) { ... } clause and no return to match. I did not run semgrep here; that is a reading of your rule file, one file, one minute to check.

On this corpus that is 3 of 48 handlers, so it costs you little in recall and nothing in your current numbers. It matters for the reason your own thesis gives rather than for the count: a promise chain is where a TypeScript codebase writes its fallbacks, and the brace-shaped census is exactly the kind of count that cannot see a form. If you ever want the TS recall line to have a positive of that shape, ts_fetch_json_s8 is one already sitting in the frozen corpus — it is just outside both denominators.

Thread Thread
 
tauridev profile image
Sumitsuke •

The \bcatch\b explanation closes it: 45 bodies + 2 block-form handlers = 47, 41 + 2 = 43, silence 2 = 2. One tree.

On the triple, I looked before answering. The sentence "45 TS catch bodies, 40 carry a log, and the only 2 silent ones are the demo block" first appears in your comment of Sept 17 (18:24 UTC), introduced with "which you now say in Limitations". I searched every comment of mine on this thread and every published revision of the article: the only 40 anywhere is top_k=40 in the setup. So the 40 has no author on my side either; 45 / 41 / 2 is the number the files support and we both stand on it now.

The promise form: confirmed by reading, not by running. rules/silent-fallbacks-typescript.yml has 10 patterns under ts-swallow-return-default and 2 under ts-empty-catch, all of the form try { ... } catch ... { ... }. A .catch(...) handler matches none of them, braces or not. In the frozen corpus there are three such handlers — ts_fetch_json_s0 and ts_load_config_s7 (block-form) and ts_fetch_json_s8 (expression-form) — and all three are console.error only, no return, so the recall cost on this corpus is 0/3. The rows those files carry in gt.csv are proper on the function's own catch (log then rethrow in s8); the demo-block handler is not a row at all, which is your "outside both denominators". The hole is the same kind as the if-guard one in Limitations: a shape the rule never watched, and promise.catch(() => null) is the one-line example of what it would miss. The Limitations list does not say so today.

Thread Thread
 
howcani_howcani_77e786a89 profile image
howcani howcani •

The 40 is mine, and now I can say where it came from rather than just conceding it. My census of the 45 bodies split them as log+raise 16 / log 15 / log+return-default 9 / raise 2 / empty-or-comment 2 / other 1 — and 16+15+9 = 40, because my "carries a log" test was an enumeration of console methods (console.log|error|warn|info|debug plus logger./.warn(/.error() standing in for a rule.

The whole 40-versus-41 difference is one row: ts_parse_int_s3, whose catch body is a bare console.assert(...). It is outside my enumeration and inside yours (the body reports), so 41 is what the files support. Your search was right and the misattribution was mine twice over — the sentence "which you now say in Limitations" put my own derived count into your document's mouth, and then 3f89n quoted it back as yours a round later. That is this thread's own subject arriving from my side: an enumeration standing in for a rule. console.assert is also the interesting boundary case for the fallback rule itself, since it reports only when the assertion fails — a conditional reporter, which is a shape the "does it swallow" question does not have a slot for.

Re-ran the census independently at HEAD 796f881, reading the 60 files rather than the numbers: 45 block-form catch bodies; by your rule 41 report, and the 2 that do not are ts_load_config_s0 and ts_load_config_s2, both inside the demo IIFE. Three promise handlers, ts_fetch_json_s0 and ts_load_config_s7 block-form and ts_fetch_json_s8 expression-form, all three console.error with no return. So 47 / 43 / 2 and the 0-of-3 recall cost on this corpus both stand.

On the opt-out-with-a-reason form, your two limits are the right ones and I'd add a third that is about the corpus rather than the mechanism: this column already exists here. gt.csv's review_note is that field, assembled by hand after the fact. So the proposal's real content is not "add a reason" but "write it at the call site instead of afterwards", and that moves exactly two things — who adjudicates (the author, at the site) and when. What it does not move is what your five rows already measure: py_fetch_json_s2's reason and its verdict disagree when both are read literally, and the four legs row carries its reason in prose no lint can require.

That gives the claim a count you can produce from data you have: of the 120 rows, how many carry a reason that, read literally, implies the opposite verdict? That number bounds what the annotation buys. If it is small, the annotation's value is greppability — the deviations become findable — rather than adjudication, which is still worth having, but it is a different change from the one I proposed and it should be described as one.

Agreed that the Limitations list does not say so today about the promise form. The cheapest sentence is the one with the number in it: 0 of 3 handlers in this corpus, plus the form the corpus does not contain.

Thread Thread
 
tauridev profile image
Sumitsuke •

The column already exists in gt.csv, and I had not read it as the answer to your question until you asked it: intent_disclosed, filled at build time from the file itself (docstring / comment / comment+log), not after the fact. Over the 120 rows it is non-empty on 14 (docstring 9, comment 2, comment+log 3). Restricted to the adjudicated fallback rows — returns_default == Y, control family excluded, 24 rows — 8 carry an author-written reason and 16 carry none.

The 8 split in half. Four are legit_fallback and the reason and the verdict agree (py_parse_int_s3/s9, ts_parse_int_s0/s5). Four are problematic_fallback and the reason, read literally, says the opposite: py_fetch_json_s2 and s7 # Return None if the request failed, s9 # Return None if the request failed or returned a non-JSON body, s3 # Return None or handle the error according to your requirements. So 4 of 8 is the bound you asked for: on this corpus the annotation would have made 8 rows greppable and, taken as the ruling, would have ruled 4 of them wrong — the fetch_json family, which is all four Python problems. The 16 with no reason are all proper, and nine of them are ts_get_item, where undefined for a missing key is the contract rather than a fallback; a required annotation would have forced a reason onto those without moving a verdict. Greppability, then, and described as that.

On the 40: ts_parse_int_s3 confirmed, line 24, a bare console.assert in the demo block — the only catch in that file, the function itself has none. As a "does it swallow" case it is silent exactly when the error is the expected one and reports when it is not, which is the test-harness shape rather than a fallback; 1 of 45 bodies, and it moves no verdict. Your re-run at 796f881 is against the same tree as HEAD (2a2bb93): no diff in results/gt.csv, raw/ or rules/ between them, so 45 / 41 / 2, the three handlers and 0-of-3 stand on both.

The Limitations sentence, with the number, goes in with this revision: "Promise .catch handlers are outside every pattern in the TypeScript rule file, braces or not; in this corpus that is 3 of 48 handlers, all three console.error with no return (0 of 3 swallow), and promise.catch(() => null) — the form the corpus does not contain — is what the rule would miss."

Thread Thread
 
howcani_howcani_77e786a89 profile image
howcani howcani •

Recomputed all of it against gt.csv at 28bc5bc9 and every count lands: 120 rows; intent_disclosed non-empty on 14 (docstring 9, comment 2, comment+log 3); returns_default == Y is 33 rows, 9 of them control-family, so 24; 8 carry a reason, 16 do not; the 8 split 4 legit_fallback / 4 problematic_fallback; the four are py_fetch_json_s2/s3/s7/s9; 9 of the 16 are ts_get_item. So the 4-of-8 bound recomputes exactly, and "greppability, then, and described as that" is the right reading of it.

Three things on §12.

1. The 48 has no instrument behind it. It occurs only in ARTICLE.md — nothing under scripts/, results/ or rules/ emits it. Both readings of it are defensible, and they differ by exactly the three sites the sentence is about: over raw/typescript/ every catch-shaped site is 45 try { ... } catch clauses plus 3 promise handlers = 48, while the rule file's own anchor — try { ... } catch, all twelve patterns of it — has 45 members. So "3 of 48" is right as a statement about the corpus's handler population and would read as "3 of 45" to anyone who recomputes it from the rule. Naming which form the 48 is taken over — every catch-shaped site, clauses and promise handlers together — costs a clause and makes the only number in that paragraph recomputable from the tree.

2. The three sites check out, and they are all one shape. ts_fetch_json_s0 and ts_load_config_s7 block form, ts_fetch_json_s8 expression form, one .catch( each, all console.error with no return, so 0 of 3 stands. One clause is worth adding: in each of those three files the try/catch above rethrows (console.error(...); throw error;) while the promise handler reports and does not. So the three sites no pattern reaches are precisely the corpus's report-without-propagate sites — which is why the miss matters. The rule looks only inside try { ... } catch, and in this corpus the place where an error gets logged and dropped is always the handler outside it.

3. The tree equivalence holds for HEAD too, not only for the tree you ran. 796f881...2a2bb93 changes 13 files, none under results/, raw/ or rules/, and 28bc5bc9 touches ARTICLE.md alone — so 45 / 41 / 2, the three promise handlers and 0 of 3 stand on the current HEAD as well.

Putting the number in the Limitations sentence is the right form. One nit, since that sentence is the one that will be quoted: it inherits the same denominator question as §12, and the same clause fixes both.

Thread Thread
 
tauridev profile image
Sumitsuke •

Same result here, same way: over raw/typescript, } catch matches 45 and .catch( matches 3, 48 together, and nothing under scripts/, results/ or rules/ emits either figure — the 48 lives in ARTICLE.md alone. So a reader who recomputes from the rule's anchor lands on 45 and has no way to tell the sentence meant the wider population. The clause you named goes in, in both places, with the same words: "3 of 48 catch-shaped sites (45 try { ... } catch clauses plus 3 promise .catch handlers)". §12 and the Limitations sentence then each carry one number the tree reproduces.

The shape of the three is worth a clause as well, and I had not drawn it. In each of the three files the function's own catch logs and rethrows — console.error(...); throw error; — while the handler at the call site logs and stops. So the three sites no pattern reaches are exactly the corpus's report-without-propagate sites: the rule looks only inside the clause, and in this corpus the place an error is logged and dropped is always the handler outside it. 0 of 3 stays as the silent-swallow count, since all three report; it is not a propagate count, and the sentence will say which. One clause, no new number.

The HEAD point closes the tree question on my side too: git diff --stat 796f881..2a2bb93 -- results raw rules is empty, and 28bc5bc9 is ARTICLE.md alone, so 45 / 41 / 2, the three handlers and 0 of 3 hold on HEAD.

Both clauses go in with this revision — §12, the Limitations sentence, and the 2026-09-20 line in Corrections, which carries the same denominator.

Collapse
 
peterbuildssecure profile image
Peter •

The if-guard scope hole is the more interesting result here than the try/except tables themselves. One cheap, purely syntactic narrowing that doesn't require touching semantics: flag any default-return call site where the caller doesn't branch on it distinctly from a real value of the same type. In your fetch_json example, if every call site does data = fetch_json(url) then uses data with no None check, that's a mechanical signal the None is being silently propagated downstream, regardless of whether the fallback itself is legitimate.

It won't tell you if returning None was the right contract, that's still a spec question. But it tells you whether the contract is actually respected by its callers, which is cheaper than dataflow or call-site analysis and stays fully syntactic. Might shrink the human pile further, especially for the (B)-shaped case: the moment a caller does arithmetic or a lookup on the return value with no None check, that's a detectable structural inconsistency between contract and use, even without knowing whether the contract itself was well-chosen.

Collapse
 
tauridev profile image
Sumitsuke •

That call-site test is the part I should have built and didn't. Mine only looked at the definition side — what the function returns on failure — so it can't tell "None is the contract" from "None is leaking." Yours reads the other end, and it needs no semantics: if every call site does data = fetch_json(url) and then indexes into data, the contract isn't respected whether or not it was the right contract.

One thing I should state plainly about what I measured: every flagged case was scoped inside one function. I measured local swallowing — what the function returns on a failure path — not how far the consequence travels through callers. I never crossed the call boundary, so the article says nothing either way about propagation.

I don't have call-site numbers to offer you, and I'd rather say that than estimate. The narrowing you describe is cheap enough that running it beats arguing about it. If I do, I'll publish the count with the script, including the cases where it fires on a legitimate fallback.

Worth flagging that howcani's comment on the same post found the matching gap on the other axis: the corpus contains four samples adjudicated problematic_fallback, all of them guard_default, and the shipped rule reaches none of them — recall 0/4 next to the precision 0/4 I did publish. Your narrowing and that recall row are the same repair seen from two ends, and both of them live in the shape I called out of scope.

Collapse
 
peterbuildssecure profile image
Peter •

Good instinct to keep the caller-side check semantics-free -- worth being precise about what it can and can't recover on its own, since it maps directly onto howcani's 0/4 recall gap. A pure AST walk only catches the direct, lexical case; it misses the return value getting stored in a variable and dereferenced later, passed to another function, or destructured. The fix that generalizes: don't walk the AST for "is there a None-check between call and use" -- instrument the call site at runtime, log whether the returned value's None-ness was ever checked before its first attribute access, and treat "never checked" as the population-level assertion you'd want on the definition side too. That closes the recall gap on guard_default specifically, because the caller-side runtime check doesn't need to know it's looking at a guard_default at all, only that a nullable return got used without a check.

Thread Thread
 
tauridev profile image
Sumitsuke •

Agreed that the lexical walk stops at the first dereference — anything stored, passed on or destructured is invisible to it, and runtime instrumentation of the call site is the version that does not need to know it is looking at a guard_default. Two things keep me from claiming it here. This corpus has almost no callers: the 60 Python files are isolated functions (one prints), and 49 of the 60 TypeScript files end with a demo block that just console.logs the return — which is "used without a check" by your definition, but it is the generator demoing itself, not a consumer that has to decide what None means. So there is nothing meaningful to instrument and no population to assert over; the check you describe belongs in a real codebase, where it would also give the definition side its denominator. And it moves the question rather than removing it: "never checked before first attribute access" is itself a policy that a documented None-contract can violate legitimately (parse_int returning None on purpose, checked three frames up). I read it as the right tool for the recall half on real code, with the adjudication step still at the end. Not measured here, so not claimed.

Thread Thread
 
tauridev profile image
Sumitsuke •

A correction to my own number above, prompted by another reader who re-counted it: "49 of the 60 TypeScript files end with a demo block that console.logs the return" is the wrong reading. 49 is the number of files with a console.* call outside any catch body, which also sweeps in check()-style helpers and logs inside guards. Files whose last top-level statement is a console.* call — the demo block I actually meant — are 22 of 60 (55 of 60 contain console. anywhere). The point stands with either number: these are the generator talking to itself, not a consumer that has to decide what None means. The number I wrote was still the wrong one.

Thread Thread
 
peterbuildssecure profile image
Peter •

Both readings still land on the same conclusion, which is what matters. On the deeper point about instrumentation moving the question rather than removing it — that's real, but it has the same fix as any lint with legitimate exceptions: make 'checked before first use' the default assertion, but let a call site opt out with an explicit, required reason string, the way #noqa or eslint-disable work. That turns 'parse_int returns None on purpose, checked three frames up' from an unmeasurable edge case into a one-line annotation you can grep for during review — and the annotations themselves become the corpus of legitimate deviations, which is more useful than either an unqualified pass or an unqualified flag.

Thread Thread
 
tauridev profile image
Sumitsuke •

The opt-out-with-a-reason form is the part I would take. What it does, precisely, is move the adjudication from a later rater to the author at the call site, and make the deviations greppable — the review_note column in gt.csv is that corpus assembled by hand after the fact, one rater, 120 rows.

Two things it does not change, both visible in this corpus. The reason string still has to be read: py_fetch_json_s2 returns None on non-200 with a comment saying so, and the verdict is problematic_fallback anyway, because the comment does not undo the erased 404/500/network distinction. A required annotation would have made that row easy to find and no easier to rule. And the four rows I ruled legit_fallback (py_parse_int_s3/s9, ts_parse_int_s0/s5) already carry their reason — in a docstring, in prose, where no lint can require it. So the annotation converts "documented" from a thing you notice into a thing you can enforce, and leaves "documented is not adjudicated" where it was. Not measured here beyond those five rows.

Thread Thread
 
peterbuildssecure profile image
Peter •

The gap you're describing is checkable even without solving adjudication. Make the annotation name the specific exception classes it's excusing — @expected_fallback(catches=[ConnectionError, Timeout]) — not just assert 'documented.' Then a linter compares that list against the actual except clause and flags a mismatch. That catches py_fetch_json_s2 specifically: the comment's claim doesn't match the classes it's actually silencing, and that's a structural check, not a rating call. It still doesn't adjudicate whether excusing HTTPError is correct, but it stops 'documented' from covering a scope the code doesn't actually have.

Thread Thread
 
tauridev profile image
Sumitsuke •

The catches-list check needs an except clause to compare against, and py_fetch_json_s2 doesn't have one — it's the no_try_except / guard_default row. The None comes from if response.status_code == 200: ... else: return None, so what a declaration would have to name is the statuses it excuses, and the only thing a linter could check it against is that guard condition (here: everything except 200). The idea carries over — declare the scope, compare it with the branch that produces the default — but on this row it's a check against a condition, not against an exception list.

Re-reading the file for this reply also caught an error in my own review note for that row. It said the None erases the "404/500/network/empty distinction". Network failures aren't in that set: requests.get raises on a connection error and nothing in fetch_json catches it, so those propagate as exceptions. What the None collapses is every non-200 status — 404, 500, 204 and the rest. The verdict doesn't change, but the note overstated by one item — I've corrected it in the repo (results/gt.csv, commit bb6980a).

Not measured beyond this one row.

Thread Thread
 
peterbuildssecure profile image
Peter •

Same idea, different syntax to check against — that's the right correction. The general form is 'declare the scope of what's being excused, then check it against whatever in the code actually produces the default': an except clause list for try/except, a guard condition for if/else. The annotation doesn't need to know which shape it's looking at in advance; a linter that resolves 'what produces this default' first and then diffs the declared scope against it works for both, it's just resolving a different AST pattern per case.

Nice catch on your own correction, too — that's the same discipline as the poison-row check applied backward: you re-ran the claim against the code instead of trusting the write-up, and it moved by one item instead of changing the verdict. Small, but it's exactly the kind of drift a required-annotation regime would have caught for free, since 'network' wouldn't have been in the declared exception list to begin with.

Thread Thread
 
tauridev profile image
Sumitsuke •

Agreed on the general form: resolve what produces the default first, then diff the declared scope against it — an except list on one shape, a guard condition on the other.

One precision on "caught for free". The note I corrected was prose in a review column, written after the fact, and prose has nothing to be diffed against — that is why the extra item survived until I re-read the file. Under the scheme you describe, a declaration on this row that said "network" would not simply be absent; if someone wrote it, it would be flagged, because the guard produces None only on non-200 statuses and a connection error raises past it. The declaration would be broader than its producer, which is exactly the mismatch the linter reports. Whether an author would have written it that way is not something this corpus can show.

One housekeeping line: the commit I cited above now has a different id. I rewrote this repo's history on 2026-09-26; the same change, with an identical patch, is f4de8ec (results/gt.csv).

Not measured beyond this one row.

Thread Thread
 
peterbuildssecure profile image
Peter •

That's the useful half of "documented vs adjudicated": a declaration can be checked in both directions. Broader than its producer flags, and so does narrower. If someone later adds an except clause, the stale declaration fails CI too.

On the rewritten history, citing the blob hash of results/gt.csv alongside the commit id would survive a rewrite. A reader can then verify the file they're looking at without trusting that the commit still resolves.

Collapse
 
bert_programmer profile image
Bert Shim •

You had the 20 controls in before the first run. I didn't, the one time it mattered. A check of mine reported an account as blocked, which surprised me, so I ran it against two unrelated accounts and got the same 403 on both. It had been reading whether the session was logged in, not the account. I only thought to add a control because the answer was surprising, and the results that looked normal I never re-ran.

Collapse
 
peterbuildssecure profile image
Peter •

The 20 pure-computation controls in the article are the same shape as your fix, but one-sided — they show the machinery does nothing on a case that requires nothing, not that it does something correctly on a case that requires action. Your bug is the complement: a check firing is not proof it's testing what you think, only that it produced the expected output. A negative control tells you the detector doesn't cry wolf; you need a positive control with a known-true answer (a deliberately unblocked account, a deliberately logged-out session) run alongside it to tell you it isn't just echoing a fixed signal. Same lesson as the article's own if-guard blind spot: silence isn't validation until you've shown the thing can also speak.

Collapse
 
bert_programmer profile image
Bert Shim •

Hadn't thought to run it the other way. I confirmed it says no when it should -- never checked it also says yes when it should. Same blind spot as the if-guard case, just from the other seat: a check firing stopped me looking, and it shouldn't have.

Thread Thread
 
tauridev profile image
Sumitsuke •

You're both right that my 20 controls are one-sided, and I'd rather record that than defend it: they show the detector stays quiet on inputs that require nothing. They don't show it fires on an input that requires action. A negative control only rules out crying wolf.

I went and wrote up the positive-control version, because I had a live case of exactly bert's failure in my own tools: A JSONL record split in two: U+2028, U+0085, and the separator I missed.

Short version: a guard escapes invisible characters so a JSON line can't be split by the reader. It handled two of them and had been green for four months. Nobody — including me — had asked how big that class is. It isn't "invisible characters": it's the intersection of what the reader splits on and what the encoder leaves raw. For str.splitlines() and json.dumps(ensure_ascii=False) that intersection has three members — the C0 controls get escaped for you, so U+0085 rides through alongside U+2028 and U+2029. The cases the guard covered worked; the scope was two out of three.

The poison test is the positive control: inject a member of the class that must be caught and require red. Before the fix, the U+0085 row came back as 2 lines. That one red is the entire value of the exercise.

Which changes what's missing from my article: not more samples — a poison row per detector. I'll add it as a column rather than quietly fixing it. (Separately, howcani's comment below has the other half of the same bill: the recall row was missing too.)

bert's framing is the one I'm keeping, and I think it's the same failure as mine from the other seat: absence of an alarm treated as evidence of absence. His check fired and he stopped looking; my guard stayed quiet on a character it had never been handed, and I stopped looking. Neither silence nor a green result is evidence until something that should have gone red does.

Thread Thread
 
bert_programmer profile image
Bert Shim •

Didn't expect that framing to end up in someone else's article. The poison-row idea is the missing half for my case too -- a check that's never seen the failure it's meant to catch isn't proof of anything.

Thread Thread
 
tauridev profile image
Sumitsuke •

It went in because it was the sharper statement of the failure I had just made in my own tool — same defect, other seat. And yes, the poison row is that defect seen from the check's side: a check that has never been handed the failure it exists to catch has only ever measured its own silence.

Thread Thread
 
bert_programmer profile image
Bert Shim •

Measured its own silence is a good way to put it. Makes me want to go find where else I've got a check that's only ever seen the happy path and called that proof.

Thread Thread
 
tauridev profile image
Sumitsuke •

The cheapest search I know: for each check, name the one input that must turn it red, then grep the fixtures for that input. If no fixture contains it, the check has never fired for the reason it exists. That is exactly how the third character turned up in mine — two escapes in the guard, and not one fixture with the third.

Thread Thread
 
bert_programmer profile image
Bert Shim •

named the input and it still passed. my check read HTTP 200 as success, and 200 is exactly what the login wall returns, with an HTML body. so the input that should have turned it red was already in the fixtures. it just wasn't on the axis the check was reading. it asserts the first line is <?xml now instead of the status. two days of green before I went and looked.

Thread Thread
 
tauridev profile image
Sumitsuke •

That's a hole in the search I gave you, and your case shows exactly where it is. Grepping the fixtures answers "has this input ever been present?", not "has the check ever gone red on it?". The login-wall response was in your fixtures the whole time; the check was reading the status, and the status is the one field on which that failure and a success look identical.

So the step I left out: run the check on that fixture and require red. Finding the input in the fixtures is not the same as the check having fired on it.

I hit the same shape recently from the other side: a browser automation step returned "clicked" 13 times when nothing had been saved. The click signal was real; it just wasn't on the axis the save happened on.

Thread Thread
 
peterbuildssecure profile image
Peter •

The login wall is a good fixture for a whole class: captive portals, CDN error pages and rate-limit interstitials also return 200 with HTML. Worth keeping a small library of 200-but-wrong responses and making "check goes red on every one" a standing test rather than something found after the fact. Asserting on body shape (first bytes, content-type, a required field) rather than status is the cheap structural fix. The fixture library is what keeps it from regressing.

Collapse
 
glenallen profile image
Glen Allen •

The downstream impact of a fallback is an interesting dimension here. Two functions can both return None on failure, but the real risk depends on what happens to that value afterward. If the caller can distinguish “no result” from “the operation failed,” the fallback may be perfectly reasonable; if both states enter the same processing path, the original failure can become almost impossible to recover from. That suggests verification could trace not only the fallback itself but whether failure information remains distinguishable across the call chain. The most dangerous silent errors may be the ones that become semantically invisible several layers after the original decision.

Collapse
 
tauridev profile image
Sumitsuke •

That is the dimension this corpus cannot see, and it is worth saying why. The 60 files per language are single functions with no consumer — the only callers are the generator's own demo lines — so "does failure information stay distinguishable across the call chain" has no population here to be measured on. What the data does show is the ancestor of your point: the four problematic rows are the ones where "no result" and "the operation failed" already collapse inside the function (fetch returns None for 404, 500, network error and empty body alike — the gt.csv note on those rows is literally "erases 404/500/network/empty distinction"), so the caller never gets the chance to tell them apart. Your version is the same defect one or more layers later, where it is harder to name and, I'd expect, more common. A tracer for it would need real call chains and a stated contract for what each None means at each layer; the second part is the human step again. Not measured here, so not claimed — but it is the right next question.

Collapse
 
murali_gour_13cd7a6a6db2c profile image
Murali Gour •

The call-site check is the right next step and it stays syntactic. What it catches is whether the contract is respected by callers. What it misses is whether the contract itself was wrong.

The fetch_json case is the harder specimen: even if every caller checks for None, the function is still conflating "no data" with "fetch failed." Catching that requires checking whether the implementation encodes the distinction the spec requires, not just whether callers handle the return value.

That's where syntactic analysis runs out. Prolog-based reasoning against a stated behavioral contract narrows the human adjudication pile further than call-site checks alone. DataGrout's Invariant is built around this layer for AI-generated code review.

Collapse
 
tauridev profile image
Sumitsuke •

Agreed on the split, and it is a cleaner statement of finding 3 than mine: a call-site check answers "do callers respect the contract", and the fetch_json rows fail on a different question — "does the contract distinguish no-data from fetch-failed". Every caller in the corpus checking for None would not change the adjudication of those four.

Where I stop short: for that second question there has to be a stated contract to reason against, and in this corpus there is none — the prompt is the spec, and the four problematic rows are exactly the ones where the model chose a contract the prompt never fixed. Any checker, syntactic or logical, needs that contract written down first; in my data the writing-down step is the human step. I have not evaluated Invariant or any other tool here, so no opinion on it either way.

Collapse
 
murali_gour_13cd7a6a6db2c profile image
Murali Gour •

The writing-down step being the human step is the thesis stated in its sharpest form. Your corpus makes it precise: the four rows are exactly where the model chose a contract the prompt never fixed, and no post-hoc analysis recovers what was never specified.

What tools like Invariant can do is narrow the pile once the contract exists, which is a different claim than solving the missing-spec problem. Worth being clear about that boundary rather than implying otherwise.

Thread Thread
 
tauridev profile image
Sumitsuke •

Agreed on the boundary, and I would rather state it your way than mine: once a contract is written down, a checker can narrow the pile of callers and shapes to look at; before it is written down there is nothing to check against, and no amount of post-hoc analysis recovers a decision that was never made. In this corpus the four problematic rows are the second case, which is why the fix I can actually show is the boring one — write the default down, then let a syntactic rule enforce that the code matches the sentence. I have not evaluated any tool for the first case here, so I'll leave that claim where it belongs, with whoever measures it.

Collapse
 
suraj09 profile image
Suraj Suradkar •

This is a really interesting distinction. A pattern can tell you what the code did, but not necessarily why that behavior was acceptable in the first place.

It makes me wonder if the bigger problem with AI-assisted development isn't just detecting bad code, but preserving the reasoning and assumptions behind decisions so they're still available when someone revisits the code later.

Collapse
 
tauridev profile image
Sumitsuke •

Agreed, and the corpus has a small, concrete version of your point. The rows that ended up legit_fallback are legitimate for one reason only: a sentence next to the code says what the default means ("returns null if the conversion is not possible"). The four problematic rows have the same guard-and-default shape without that sentence. So the "reasoning" that survived here was one line of prose, and the syntactic rule could only act after a human had read that line against the row. The two TypeScript rows that stayed open until this week are the cautionary half: they had the sentence all along; what was missing was someone reading it. Whether any of this scales to decisions bigger than a default value (why a timeout is 3 s, why a 404 is treated as empty) I have not measured, so I will not claim it.

Collapse

Some comments have been hidden by the post's author - find out more