DEV Community

Sneha M K
Sneha M K

Posted on

Claude Can't Say "Done" Until the Code Is Safe: A Security Verification Loop for Claude Code

I asked Claude Code to add an AI chat feature. It worked. It also:

  • hardcoded my API key,
  • rendered the model's reply with innerHTML,
  • and ran eval() on it.

Then it said, "All done!" 🙃

Anthropic's Claude Code team recently wrote about verification loops: Claude checks its own work and loops back to fix problems before responding. Most examples verify that tests pass. Nobody was verifying that the code was safe. So I built that.

How it works

Claude Code hooks run your scripts at fixed moments. The Stop hook fires when Claude is about to finish. If the hook exits with code 2, Claude isn't allowed to stop, and whatever you print to stderr is sent back to it as feedback.

That's the whole loop:

Claude says done → hook scans the diff → exit 2 with findings → Claude fixes → hook passes → done.

Step 1: register the hook

.claude/settings.json

{
  "hooks": {
    "Stop": [{
      "hooks": [{
        "type": "command",
        "command": "node \"$CLAUDE_PROJECT_DIR\"/.claude/hooks/security-verify.mjs"
      }]
    }]
  }
}
Enter fullscreen mode Exit fullscreen mode

Step 2: the rules

The hook scans only the lines Claude changed (git diff HEAD -U0 plus new files), so old code never blocks you. Five rules, each aimed at a mistake AI code actually makes:

const RULES = [
  { id: 'hardcoded-secret',        test: l => /sk-(ant-|proj-)?[\w-]{20,}|AKIA[0-9A-Z]{16}/.test(l) },
  { id: 'secret-in-client-bundle', test: l => /(NEXT_PUBLIC_|VITE_)\w*(KEY|SECRET)|dangerouslyAllowBrowser:\s*true/.test(l) },
  { id: 'llm-output-as-html',      test: l => /innerHTML|dangerouslySetInnerHTML/.test(l) && /\b(reply|response|completion)/i.test(l) },
  { id: 'unsafe-html-sink',        test: l => isDynamicHtml(l) }, // static strings are allowed
  { id: 'dynamic-code-exec',       test: l => /\beval\s*\(|new\s+Function\s*\(/.test(l) },
];
Enter fullscreen mode Exit fullscreen mode

llm-output-as-html is the one I care most about. Suppose a prompt injection gets into your model's reply; innerHTML = replyhands the attacker XSS. Model output is untrusted input, always.

Step 3: block, with a way out

const findings = scan(changedLines());
if (findings.length === 0) process.exit(0);          // clean: Claude may finish

if (attempts > MAX_ATTEMPTS) {                        // stuck: hand over to the human
  writeFileSync('.claude/security-report.md', report);
  process.exit(0);
}
console.error(`security-verify blocked completion:\n${report}`);
process.exit(2);                                      // Claude must keep working
Enter fullscreen mode Exit fullscreen mode

The attempt cap matters. A Stop hook that always exits 2 can loop forever. I count attempts per session and, after 3 tries, stop blocking and write a report for a human instead.

Step 4: teach Claude how to fix it

The hook says what is wrong. A skill in .claude/skills/security-verify/SKILL.md says how to fix it: move keys behind a server route, use textContent or DOMPurify.sanitize(), use JSON.parse instead of eval. It also tells Claude not to just rephrase code to dodge the regex.

False positive? Claude can add // verify-ignore: on that line, and it has to tell you why. The opt-out is visible in review, not silent.

The result

On my test repo, the first "done" was blocked with 5 findings. Claude moved the API call to a server route, switched to textContent and JSON.parse, the hook passed, and its final message told me to rotate the key that had already appeared in the diff.

That last part is the real win: I didn't have to remember to check.

Honest limits

These are pattern rules, not a full security audit. They catch the common AI mistakes cheaply on every single turn. Keep your real SAST, code review, and secret scanning in CI too.

📦 Get the drop-in .claudefolder (hook + settings + skill, zero dependencies):

Top comments (7)

Collapse
 
elijahbrown profile image
Elijah Brown •

Nice loop. A sixth rule I'd add for AI-written forms: flag a new route or server action that writes an email or phone field without a server-side schema parse, because it's easy for generated code to keep the validation only in the client component while the endpoint accepts anything.

Collapse
 
thesnehamk profile image
Sneha M K •

Your postmortem on the test gate that passed a failing build was in the back of my mind while writing this. The whole loop depends on exit code 2 reaching Claude Code, so if someone wraps the hook in a pipe like node security-verify.mjs | tee log, the gate silently stops working. I should probably add that as a warning in the post.

Curious: have you seen Stop hooks cause loops in practice, or does a retry cap like mine cover it?

Collapse
 
elijahbrown profile image
Elijah Brown •

A retry cap like yours is the safer default; the loop risk usually comes from a hook re-invoking itself or from retries resetting the same condition. I would log a correlation ID and cap by both attempts and elapsed time.

Collapse
 
elijahbrown profile image
Elijah Brown •

I have seen loops come from a hook treating every nonzero exit as retryable, especially when the verifier is invoked again by the retry path. A small retry cap is a good backstop, but I would also stop when the same input and failure recur. The pipe case is a separate failure mode, so documenting the exit-code warning is worthwhile.

Collapse
 
elijahbrown profile image
Elijah Brown •

That is a useful failure mode to call out. I have not seen Stop hooks loop in practice, but a retry cap is the right guard; I would also surface the exit code so the failure cannot look like a clean pass.

Collapse
 
rulestack profile image
Rulestack •

Your test repo got real fixes, so I'm curious: have you seen Claude reword a line just enough to slip past one of the five rules, despite the skill telling it not to, or has that instruction held so far?

Collapse
 
thesnehamk profile image
Sneha M K •

Great question, and the honest answer is: not yet, but my sample is small, so I wouldn't call it "held."

And it would be trivially easy. el['inner' + 'HTML'] = reply or window'ev' + 'al' sail straight past the regexes. The "don't reword to dodge" line in the skill is a soft control. It's the model following an instruction, not the gate enforcing anything. There is some layering: moving reply into another variable dodges llm-output-as-html, but unsafe-html-sink still catches the dynamic innerHTML. Computed property access beats both, though.

Two things I'm considering:

Flag computed access to sink names (['inner…, ['ev…) as its own rule, since legit code rarely builds those names dynamically.
Swap regex for an AST pass (acorn, or eslint-plugin-no-unsanitized) so it doesn't matter how the line is spelled.

I'll also start logging every blocked line next to Claude's fix, so "did it reword or really fix it" becomes measurable instead of a vibe. You've run agents in public far longer than I have: have you caught yours satisfying a check's letter instead of its intent?