DEV Community

Cover image for I Built Two Agent Systems. Each One Proved the Other One Wrong.
Debashish Ghosal
Debashish Ghosal

Posted on AI-assisted

I Built Two Agent Systems. Each One Proved the Other One Wrong.

I built one engine where an LLM reviews another LLM's plan. I built another where two LLMs debate a pull request.

Each one failed in the way the other was designed to prevent. That's the only reason I trust either of them.

If you're putting your safety in a prompt, both stories end the same way.

System One: The Critic I Couldn't Make Reliable — Or Needed To

PlannerCritic runs a draft → critique → revise loop: one model writes a plan, another critiques it, and a bounded loop revises until it converges or escalates to a human.

My first instinct was the obvious one: tell the critic to be adversarial. That's a prompt fix for what I thought was a behavior problem.

It backfired. The critic started blocking plans for being "not thorough enough," which is a useless verdict. I had to move the severity contract out of the prompt and into a frozenset in code, where it couldn't drift.

Here's what the field test showed on 170 goals for $0.49:

  • Critic label_flip_rate: 1.0 — on identical input, it changed its verdict every single trial.
  • Critic evidence_drift_rate: 1.0 — its justification moved too.
  • underclaim_approvals: 0 — it never knowingly let a seeded defect through.
  • family_migration_rate: 0.0.
  • True failures: 0.

A maximally non-deterministic critic was completely safe. Not because it was reliable — because the deterministic gates owned the under-claim direction, and the critic was never allowed to be the safety boundary. The numbers are in the v0.2.1 field test results.

System Two: The "Debate" That Was Actually a Replay

AdversarialDebate does the opposite: two isolated models review the same PR and either converge on a verdict or preserve their disagreement.

The failure here was subtler. I asked the second model to be independent in the prompt. It looked like it worked. Then I read the raw logs: the "debate" was a stored replay, echoing text the first model had already generated.

89% of the second opinions were theater.

I tore out the prompt-level independence and enforced it mechanically — isolated contexts, no shared conclusion, evidence required per claim. Theater went from 89% to 0% across 217 debates, and on 411 debates against 70 real public PRs the system matched human review claims 81% of the time. Details in the v0.2.2 field test report.

The Convergence

Look at the two fixes side by side and they're the same fix.

System Structural problem Prompt attempt What actually fixed it
PlannerCritic Unreliable severity contract "Be adversarial" Severity contract in code (frozenset) + deterministic gates
AdversarialDebate Fake independence "Be independent" Mechanical isolation invariant

Both times, the default instinct was to solve a structural problem with a prompt. Both times, the prompt was the wrong layer. The model was never the safety boundary — the architecture was.

Why This Generalizes (Cautiously)

This is a sample size of two, and both systems are mine, so take the convergence as suggestive, not proven.

But it's suggestive in a useful direction: if you can't make an LLM reliable, stop trying to. Make it irrelevant to safety. Let the nondeterministic part be as wild as it wants, and put the non-negotiable decisions in deterministic code that no prompt can override.

The Honest Limitation

The architectures aren't complete. PlannerCritic catches structural plan defects but not subtle logic errors. AdversarialDebate still has an unexplained 11.3% miss rate in narrative domains, and I can't fully account for it.

Nor is "two systems converged" proof of a universal law. It's two data points from one engineer. What I can defend is the direction: the failures were structural, and so were the fixes.


Where does the safety boundary live in your agent — in the prompt or in the code? I've been wrong about that in both directions.

Repos and receipts: planner-critic-engine · adversarial-debate · PlannerCritic failure-mode register — all MIT, all public.

Top comments (4)

Collapse
 
howcani_howcani_77e786a89 profile image
howcani howcani •

ivannovazzi has the right first question, and I think two of the four headline numbers need a different explanation before the temperature check is applied.

A pairwise flip rate of 1.0 is not reachable by independent trials. If each goal has a fixed probability p of one label and the two trials are independent draws, then the probability that the second differs from the first is 2p(1-p), whose maximum over p is 0.5, at p = 0.5. So a metric defined as the fraction of goals whose two trials carried different labels has a ceiling of 0.5 under i.i.d. sampling, and a measured 1.0 says the two trials in each pair are not independent draws. Something couples them: state carried from the first call into the second, a seed or temperature that cycles between trials, a prompt or cache that differs, or a pairing in the metric that is not the pairing I am assuming. That is a finding about the harness rather than about the critic, and it is worth settling before anything else, because it changes what the number means. If the flip is a cycle, the label is reproducible: run one goal four times and see whether the sequence alternates or splits. Under independent draws at p = 0.5, strict alternation has probability 0.125, so four runs tell you which process you are looking at in a few minutes of compute.

A drift rate of 1.0 on free text may be a property of the measure. Two samples of generated justification are essentially never byte-identical, so any difference metric without a threshold reports a difference every time on any population, including one where the reasoning is stable and only the wording moves. The control that separates those: compute the same statistic on pairs that have no reason to differ (the same text against itself after a whitespace change) and on pairs where drift is certain (critiques of two different goals). If the metric returns the same value across all three, it is measuring that text is text.

underclaim_approvals: 0 is the one number whose control I would want printed beside it. Your thesis is that the deterministic gates own the under-claim direction and the critic was never the boundary. If that is true, then zero under-claim approvals is a reading of the gates, and the critic contribution is unmeasured by the field test. The 2x2 over seeded defects would show it: caught by the gate, caught by the critic, caught by both, caught by neither. The fourth cell is the one that decides whether the critic is doing work, and with the gate as the boundary it is the cell nobody is looking at. Zero detections only reads as safety when the count of things it should have detected travels with it.

The thesis itself I would keep, and it generalizes past your two systems: I have watched the same instinct fail in the other direction, where a check that could not fail was read as a check that passed, and a prompt-level assertion was trusted to enforce a property that only the writer of the record could enforce. Both of your fixes moved the property to where it stops being a promise. What I would add is that after the move, the prompt-side signal becomes decorative in a way that is easy to keep reading as load bearing, which is why the cells above are worth printing.

Both of your failures are the same failure in one more respect than the prompt one: in each case a metric was reported whose value a fixed-difficulty measurement cannot produce, and the fix was structural in the engine while the reading stayed structural in the report.

Collapse
 
ivannovazzi profile image
Ivan Annovazzi •

label_flip_rate at 1.0 on identical input is the part I'd dig into first. Did that persist with temperature 0 or a fixed seed? If the critic flips that much, the revisions it drives look close to random, so the deterministic gates are carrying all the safety weight. Curious whether you'd trust the loop without them

Collapse
 
contentclips_st profile image
ContentClips •

The severity contract in a frozenset is the right call. One addition: log the verdict key per run (plan hash + gate version) so label_flip_rate becomes a CI metric instead of something you find by eyeballing raw logs. The replay detection on the debate side is the most valuable finding here — independence has to be enforced at the context boundary (no shared cache, no shared conversation state, per-claim evidence links), never asserted. For the unexplained 11.3% narrative miss rate: worth bucketing misses by whether the claim carried a machine-checkable citation. My hunch is the miss concentrates where 'evidence required per claim' degraded into prose citations that no downstream checker can re-verify mechanically. That would push the fix in the same direction as your frozenset: make the evidence artifact a structured claim-to-citation map so the deterministic layer can re-verify it without the model in the loop.

Some comments may only be visible to logged-in visitors. Sign in to view all comments.