DEV Community

Juber Shaikh
Juber Shaikh

Posted on

I stopped trusting my AI agents' "done" — here's the boring gate that fixed it

I use AI coding agents every day for real client work. A few months ago I noticed a pattern, and once you see it you can't unsee it:

The agent says "done". I push. Tests fail. Or lint breaks. Or the same wrong assumption comes back next session, because nothing was ever written down anywhere.

Each instance was easy to fix. But the pattern kept costing me. And the pattern wasn't a model problem — it was a verification problem. Nobody was checking the work before it left the machine.

The dumb fix that worked

I stopped trying to make the agent smarter and built a gate instead. The gate is deliberately dumb — no intelligence in it at all. It's a Node CLI called forgekit (I built it, it's MIT, zero runtime dependencies):

1. forge verify — make "done" mean something checkable

It detects the repo's test suites, runs them, and reports the truth: PASS, FAIL, INCOMPLETE, or PARTIAL. The key decision: a run that only covered part of the repo is reported as partial — never as green. Most of my "done but broken" failures died right here.

forgekit verify

2. forge precommit — check before it lands

Looks at what's actually staged and scans for secrets before anything pushes. The unglamorous stuff that saves you exactly once, and that once is worth it.

forgekit precommit

3. forge impact — "if I touch this, what breaks?"

A blast-radius estimate before an edit. Honest caveat, and I put it in the README too: it's a heuristic code graph, not a sound call graph. It misses files as well as over-warns. I treat its output as advisory. It's the part I trust least — real repos are the only way to harden a heuristic.

4. forge ledger / forge remember — memory with receipts

Lessons and decisions live in the repo as claims, each carrying its evidence. Only tests, CI, or a human raise a claim's confidence. So the agent stops re-learning the same thing every session — and when it "remembers" something, you can see why it believes it.

forgekit ledger

The workflow rule

No push without forge verify + forge precommit green. It's a rule, not a suggestion. The repeated-failure class that motivated the whole project is gone from my workflow.

What I deliberately didn't build

A sandbox — guardrails reduce risk, they're not a security boundary. And anything claiming formal verification: "proof-carrying memory" is the feature's name, not a theorem. When you're selling reliability, naming things honestly is part of the product.

Try it

npm install -g @codewithjuber/forgekit → forge init. One init emits native config for Claude Code, Codex, Cursor, Gemini, Aider, Copilot, Windsurf, Zed, Continue and OpenClaw. (Claude Code is what I use daily, so that's the deepest-tested integration.)

Repo: https://github.com/CodeWithJuber/forgekit

If you try it, tell me where the impact analysis misses. That's the backlog I'm building against.

Top comments (8)

Collapse
 
ritu_kumari_45352a2786b76 profile image
Ritu Kumari •

Good read. I like that partial verification doesn’t count as a pass. Running a few checks isn’t the same as checking the whole change. I’d rather know what the agent skipped than get told everything’s done and find the gaps while reviewing the PR.

Collapse
 
zubair_shaikh_a22cea567b0 profile image
Juber Shaikh •

Exactly. That’s the main reason I wanted PARTIAL to be explicit instead of treating it as “good enough.” If something didn’t run, I’d rather surface that immediately than let a green-looking result create false confidence during review.

Collapse
 
codxzaheer profile image
Zaheer khan •

Completely agree with your philosophy here.

Collapse
 
zubair_shaikh_a22cea567b0 profile image
Juber Shaikh •

Thanks, glad it resonated. For me, evidence matters more than the agent saying it’s done.

Collapse
 
nikhil_darji_bed2fc186657 profile image
Nikhil Darji •

How do you handle flaky tests here? Would be useful to distinguish an unrelated failure from something the agent actually broke, without letting it dismiss every failure as unrelated.

Collapse
 
zubair_shaikh_a22cea567b0 profile image
Juber Shaikh •

Good question. I wouldn’t let the agent call a failure flaky on its own. I’d compare the same check against the base branch and keep the retry results visible. One passing rerun wouldn’t be enough for me to dismiss it.

Collapse
 
aman_shekh_521987c21f181e profile image
Aman shekh •

Curious how you decide which checks are enough for a change. A small diff in shared code can affect a lot more than a large diff in an isolated component.

Collapse
 
zubair_shaikh_a22cea567b0 profile image
Juber Shaikh •

Exactly. I think the checks should depend more on impact than diff size. Shared or critical code needs broader verification, even if the change itself is tiny.