DEV Community

Ali Suleyman TOPUZ
Ali Suleyman TOPUZ

Posted on Originally published at topuzas.medium.com on

Multi-Agent Systems Don’t Have a Consistency Problem. They Have a Database Problem.

The bugs I kept blaming on the model turned out to be the same bugs I was fixing in 2012. The fix was boring. It worked.

The night one customer got refunded twice

The order was number 4471. It was a boring order: two items, one of them out of stock, a customer who wanted a partial refund. I had a small multi-agent setup running against a sandbox: an order agent that owned order state, and a support agent that talked to the customer and could trigger refunds through a tool.

Here’s what happened, reconstructed from the logs the next morning.

The support agent decided a partial refund was right and called the issue_refund tool. The payment sandbox was slow that evening. The tool call hit a 30 second timeout. The orchestrator did exactly what I had told it to do on timeouts: it retried the step. The model, seeing the same conversation, made the same decision and called issue_refund again.

Both calls went through. The first one had not failed. It had just been slow.

Meanwhile, the order agent had read order 4471 at status = 'awaiting_stock', decided the missing item should be cancelled, and wrote status = 'partially_cancelled'. A few hundred milliseconds later, the support agent, working from its own earlier read of the same row, wrote status = 'refund_pending'. Last write won. The cancellation vanished from the record. Nobody raised an error, because from the database's point of view nothing was wrong. Two clients wrote to a row. That's what rows are for.

So by morning I had:

  • a customer refunded twice for one item
  • an order whose status said “refund pending” while the warehouse thought it was cancelled
  • a Slack thread where I had typed “the agents are being inconsistent again, need a better system prompt”

That last line is the one I’m embarrassed about. Because I spent the next two evenings rewriting prompts. I added “never issue a refund twice” to the support agent. I added “check the current status before writing” to the order agent. It got slightly better and then broke again the next time the sandbox was slow.

Nothing about this failure was an AI problem. It was a retry without an idempotency key, and a lost update caused by two writers with stale reads. I have fixed both of those bugs in systems with zero LLMs in them. The model was just the part of the pipeline I was most suspicious of, so it’s where I looked first.

The LLM is not the bug. The LLM is a very slow, very creative client.

This is the argument of the whole article, so let me state it plainly.

Most “multi-agent consistency” failures are classic distributed-systems failures. Race conditions, lost writes, duplicated side effects, out-of-order messages. They show up in agent systems because agent systems are distributed systems: multiple independent actors, each with a partial and possibly stale view of shared state, communicating over unreliable channels, taking actions with real side effects.

An LLM in the loop changes a few things about the shape of these problems, and all of them make things worse:

  • Latency is huge and variable. A model call can take 800ms or 40 seconds. The window between “read state” and “write decision” is enormous compared to a normal service, so races that would be rare in a CRUD API become routine.
  • Retries are natural. Timeouts, rate limits, and flaky tool calls mean every framework retries. Retrying a non-idempotent side effect is how you refund someone twice.
  • Decisions are not replayable. If you retry a step, the model may decide something slightly different the second time. So you can’t even assume the retry is a duplicate of the original. It might be a new, conflicting intent.

None of this is fixed by a smarter prompt. You can’t prompt your way out of a lost update any more than you can fix a race condition in Java by adding a comment that says “please don’t race.”

I found JIN’s piece Multi-Agent Consistency Isn’t an AI Problem after I’d already gone through this, and it put words to something I’d felt but not articulated: the moment you have more than one writer, you’ve signed up for distributed-systems discipline whether you wanted to or not.

Every orchestration framework hits the same three walls

I’ve built small things with LangGraph, poked at CrewAI, used the OpenAI Agents SDK’s handoffs, and read through how Microsoft Agent Framework does checkpointing. They’re all useful. They solve real problems around control flow, state passing, and resuming after crashes.

But none of them solve the three problems that actually bit me, because those problems live below the orchestrator, at the boundary where an agent touches the outside world.

+-------------------------------+-----------------------------+---------------------------------------+
| Problem | What orchestrators give you | What they usually DON'T give you |
+-------------------------------+-----------------------------+---------------------------------------+
| Idempotency | Step retries, checkpoints, | A guarantee that the external side |
| (same intent, one effect) | resume from last state | effect (payment, email, booking) |
| | | fires exactly once across retries |
+-------------------------------+-----------------------------+---------------------------------------+
| Ordering | A graph or handoff sequence | Protection against two branches or |
| (writes land in a sane order) | inside ONE workflow run | two runs writing the same record with |
| | | stale reads |
+-------------------------------+-----------------------------+---------------------------------------+
| Conflict resolution | "Last node wins" state | A record of WHO asserted WHAT based |
| (two agents disagree) | merging, reducers | on WHICH version, so disagreements |
| | | can be reconciled, not overwritten |
+-------------------------------+-----------------------------+---------------------------------------+
Enter fullscreen mode Exit fullscreen mode

Checkpoints are the one that fooled me longest. Durable execution (Shrashti Singhal’s write-up on agents that survive crashes and restarts is a good intro) is genuinely great. But a checkpoint tells you where your workflow was. It does not tell you whether the payment API received your request before the process died. Resuming from a checkpoint and re-running the “issue refund” step is exactly the double-refund scenario, just triggered by a crash instead of a timeout.

The orchestrator serializes steps inside its own view of the world. Your database, your payment provider, and your calendar API don’t live inside that view.

What 16 agents writing a C compiler actually needed

If you want to see what this looks like at real scale, Anthropic published a write-up earlier this year on building a C compiler with a team of parallel Claude agents. Sixteen agents, roughly 2,000 Claude Code sessions, around $20K in API cost, and about 100,000 lines of Rust at the end, capable of compiling a real Linux kernel.

What struck me reading it wasn’t the model capability. It was how much of the engineering was coordination plumbing, and how old-fashioned that plumbing was.

Each agent ran in its own container with its own clone of a shared git repository. To take a task, an agent wrote a lock file into a shared directory and pushed it. If another agent had already claimed that task, git’s push rejection forced the second agent to pick something else. Merge conflicts were constant and the agents resolved them the normal way.

Read that again with a database hat on. That’s a claim table with compare-and-swap semantics , implemented with git as the storage engine. The push either succeeds against the version you based it on, or it’s rejected and you re-read. That’s optimistic concurrency control.

And the failure they described is also a textbook one: when every agent ran into the same big problem (the kernel build), they all tried to fix the same bug and kept overwriting each other’s work. More parallelism produced less progress, because there was no way to partition the work so that writes didn’t collide. The fix was to change how the work was split so agents stopped touching the same state, not to make the agents smarter.

That’s the lesson I took from it. At 16 agents, the discipline that kept things consistent was locks, versioned state, and partitioned work. The models were the easy part.

Pattern 1: Idempotency keys on every agent-triggered side effect

The rule I follow now: no agent tool that causes an external side effect runs without an idempotency key, and the key is never generated by the model.

That second half matters. If you let the LLM invent a request ID, a retry gets a fresh ID and your dedupe does nothing. The key has to be derived deterministically from the intent: which workflow run, which step, which target, which payload.

Here’s the table:

CREATE TABLE side_effects (
    idempotency_key TEXT PRIMARY KEY,
    effect_type TEXT NOT NULL,
    request_hash TEXT NOT NULL,
    status TEXT NOT NULL DEFAULT 'pending', -- pending | done | failed
    result JSONB,
    created_at TIMESTAMPTZ NOT NULL DEFAULT now(),
    completed_at TIMESTAMPTZ
);
Enter fullscreen mode Exit fullscreen mode

And the wrapper every side-effecting tool goes through:

import hashlib
import json

import psycopg
DSN = "postgresql://agents:agents@localhost:5432/agents"

def _hash(obj) -> str:
    return hashlib.sha256(json.dumps(obj, sort_keys=True).encode()).hexdigest()

def effect_key(run_id: str, step: str, target: str, payload: dict) -> str:
    """Deterministic key: same intent => same key, across retries and restarts."""
    return _hash({"run": run_id, "step": step, "target": target, "payload": payload})

class EffectInProgress(Exception):
    pass

def run_once(conn, key: str, effect_type: str, payload: dict, do_effect):
    req_hash = _hash(payload)
    claimed = conn.execute(
        """
        INSERT INTO side_effects (idempotency_key, effect_type, request_hash)
        VALUES (%s, %s, %s)
        ON CONFLICT (idempotency_key) DO NOTHING
        RETURNING idempotency_key
        """,
        (key, effect_type, req_hash),
    ).fetchone()
    if claimed is None:
        status, result, existing_hash = conn.execute(
            "SELECT status, result, request_hash FROM side_effects WHERE idempotency_key = %s",
            (key,),
        ).fetchone()
        if existing_hash != req_hash:
            raise ValueError(f"Key {key} reused with a different payload")
        if status == "done":
            return result # duplicate call: return the original outcome, do nothing
        raise EffectInProgress(key) # someone else is mid-flight; caller should back off
    try:
        # Pass the same key downstream so the provider can dedupe too.
        result = do_effect(payload, idempotency_key=key)
    except Exception:
        conn.execute(
            "UPDATE side_effects SET status = 'failed' WHERE idempotency_key = %s", (key,)
        )
        raise
    conn.execute(
        """
        UPDATE side_effects
        SET status = 'done', result = %s, completed_at = now()
        WHERE idempotency_key = %s
        """,
        (json.dumps(result), key),
    )
    return result

# usage
with psycopg.connect(DSN, autocommit=True) as conn:
    payload = {"order_id": 4471, "amount_cents": 1899, "reason": "out_of_stock"}
    key = effect_key(run_id="run-8f2c", step="issue_refund", target="order:4471", payload=payload)
    run_once(conn, key, "refund", payload, do_effect=lambda p, idempotency_key: {"refund_id": "rf_123"})
Enter fullscreen mode Exit fullscreen mode

An honest caveat, because I got this wrong the first time: there’s a window where the effect succeeded but the process died before writing done. The row stays pending forever. Two things cover it. First, pass the same key to the downstream provider (Stripe and most payment APIs accept one), so even if you retry, they dedupe. Second, run a small reconciliation job that looks at old pending rows and asks the provider what actually happened. This is not glamorous. It's also the difference between "we refunded twice" and "we didn't."

A note on failed: I deliberately don't auto-retry failed rows with the same key. A failure might mean the model's intent was wrong, and I'd rather the next attempt go back through the agent's decision with fresh state.

Pattern 2: Optimistic locking in SQL, not in the orchestrator

The lost update on order 4471 happened because two agents read the same row, thought for a while, and wrote blindly. The orchestrator couldn’t prevent it because the two agents were in different workflow runs. The orchestrator had no idea they were touching the same record.

The database did know. I just hadn’t asked it to care.

Add a version column and make every write a compare-and-swap:

ALTER TABLE orders ADD COLUMN version INT NOT NULL DEFAULT 0;
-- Every agent write looks like this. Zero rows updated = someone else got there first.
UPDATE orders
SET status = %(new_status)s,
       version = version + 1
WHERE id = %(order_id)s
  AND version = %(expected_version)s
RETURNING version;
Enter fullscreen mode Exit fullscreen mode

The important design decision is what happens on conflict. My first instinct was to just retry the write. That’s wrong. The agent’s decision was based on state that no longer exists. The correct move is to re-read and re-decide :

class ConcurrencyConflict(Exception):
    pass
def agent_update_order(conn, order_id: int, decide, max_attempts: int = 3):
    for attempt in range(1, max_attempts + 1):
        order = conn.execute(
            "SELECT id, status, version, items FROM orders WHERE id = %s", (order_id,)
        ).fetchone()
        # The slow part: an LLM call. Other agents can write during this window.
        decision = decide(order)
        if decision["new_status"] == order[1]:
            return order # nothing to change
        row = conn.execute(
            """
            UPDATE orders SET status = %s, version = version + 1
            WHERE id = %s AND version = %s
            RETURNING version
            """,
            (decision["new_status"], order_id, order[2]),
        ).fetchone()
        if row is not None:
            return row
        # Lost the race. Loop back, re-read fresh state, and let the agent decide again.
    raise ConcurrencyConflict(f"order {order_id}: gave up after {max_attempts} attempts")
Enter fullscreen mode Exit fullscreen mode

Bounded attempts matter. If two agents keep flipping a record back and forth, that’s not a race, that’s a genuine disagreement, and retrying forever just burns tokens. That’s what Pattern 3 is for.

Since a lot of my day job is .NET, here’s the same thing with Npgsql:

using Npgsql;
public sealed class ConcurrencyConflictException(string message) : Exception(message);
public static async Task<int> UpdateOrderStatusAsync(
    NpgsqlConnection conn, long orderId, string newStatus, int expectedVersion)
{
    const string sql = """
        UPDATE orders
        SET status = @status, version = version + 1
        WHERE id = @id AND version = @expected
        RETURNING version
        """;
    await using var cmd = new NpgsqlCommand(sql, conn);
    cmd.Parameters.AddWithValue("status", newStatus);
    cmd.Parameters.AddWithValue("id", orderId);
    cmd.Parameters.AddWithValue("expected", expectedVersion);
    var result = await cmd.ExecuteScalarAsync();
    if (result is null)
        throw new ConcurrencyConflictException(
            $"Order {orderId} changed since version {expectedVersion}. Re-read and re-decide.");
    return (int)result;
}
Enter fullscreen mode Exit fullscreen mode

If you’re on EF Core, a [ConcurrencyCheck] or row-version property gives you the same behavior and throws DbUpdateConcurrencyException. Same idea, same "re-read and re-decide" handling.

Pattern 3: A claim graph, so disagreements stop being silent

Optimistic locking tells you that two agents collided. It doesn’t tell you why, and it doesn’t keep the losing side’s reasoning. In my double-refund case, the order agent’s “partially cancelled” conclusion was actually correct, and it disappeared without a trace.

So I stopped letting agents write conclusions directly into business tables. Agents write claims. A claim says: this agent, in this run, asserted this value for this fact, based on this version of the world, for this reason.

CREATE TABLE claims (
    id BIGSERIAL PRIMARY KEY,
    subject TEXT NOT NULL, -- e.g. 'order:4471'
    predicate TEXT NOT NULL, -- e.g. 'status'
    value JSONB NOT NULL,
    agent_id TEXT NOT NULL,
    run_id TEXT NOT NULL,
    based_on_version INT, -- version of the subject the agent read
    evidence JSONB, -- tool outputs / reasoning summary
    supersedes BIGINT REFERENCES claims(id),
    asserted_at TIMESTAMPTZ NOT NULL DEFAULT now()
);

CREATE INDEX claims_subject_idx ON claims (subject, predicate, asserted_at DESC);
Enter fullscreen mode Exit fullscreen mode

It’s a graph because of supersedes: when an agent updates its view after seeing another claim, it links to the claim it's replacing. Two claims on the same fact, based on the same version, neither superseding the other, with different values: that's a conflict. You can find those with plain SQL:

SELECT a.id AS claim_a, a.agent_id AS agent_a, a.value AS value_a,
       b.id AS claim_b, b.agent_id AS agent_b, b.value AS value_b
FROM claims a
JOIN claims b
  ON a.subject = b.subject
 AND a.predicate = b.predicate
 AND a.id < b.id
WHERE a.subject = %(subject)s
  AND a.value <> b.value
  AND a.based_on_version = b.based_on_version
  AND b.supersedes IS DISTINCT FROM a.id
  AND NOT EXISTS (SELECT 1 FROM claims c WHERE c.supersedes IN (a.id, b.id));
Enter fullscreen mode Exit fullscreen mode

Then a reconciler decides what gets committed to the real orders table. I start with deterministic rules and only escalate when rules can't decide:

# Which agent is authoritative for which predicate. Boring on purpose.
AUTHORITY = {
    "status": ["order_agent", "support_agent"], # order agent wins on status
    "refund_amount": ["support_agent"],
}

def reconcile(conn, subject: str, predicate: str):
    rows = conn.execute(
        """
        SELECT id, agent_id, value, evidence FROM claims
        WHERE subject = %s AND predicate = %s
          AND id NOT IN (SELECT supersedes FROM claims WHERE supersedes IS NOT NULL)
        ORDER BY asserted_at DESC
        """,
        (subject, predicate),
    ).fetchall()
    distinct_values = {json.dumps(r[2], sort_keys=True) for r in rows}
    if len(distinct_values) <= 1:
        return ("agreed", rows[0] if rows else None)
    ranking = AUTHORITY.get(predicate, [])
    ranked = sorted(
        rows, key=lambda r: ranking.index(r[1]) if r[1] in ranking else len(ranking)
    )
    if ranked[0][1] in ranking:
        return ("resolved_by_authority", ranked[0])
    # No rule applies: park it for a human (or a judge agent with all evidence attached).
    return ("needs_review", rows)
Enter fullscreen mode Exit fullscreen mode

Is this more work than letting agents write directly? Yes. But the first time a conflict landed in my review queue with both agents’ evidence side by side, I understood in ten seconds what had happened. Compare that to two evenings of rewriting prompts.

Running all of this locally (no paid API needed)

You don’t need a hosted model to reproduce any of this. The races happen at the database layer. I run Postgres and Ollama in Docker and use a small local model as the “agent brain.”

# docker-compose.yml
services:
  postgres:
    image: postgres:16
    environment:
      POSTGRES_USER: agents
      POSTGRES_PASSWORD: agents
      POSTGRES_DB: agents
    ports:
      - "5432:5432"
ollama:
    image: ollama/ollama:latest
    ports:
      - "11434:11434"
    volumes:
      - ollama:/root/.ollama
volumes:
  ollama:

docker compose up -d
docker compose exec ollama ollama pull qwen2.5:7b
pip install "psycopg[binary]" requests pytest
Enter fullscreen mode Exit fullscreen mode

A decision function that uses the local model and forces JSON output:

import json
import requests
OLLAMA_URL = "http://localhost:11434/api/chat"

def ollama_decide(order) -> dict:
    order_id, status, version, items = order
    resp = requests.post(
        OLLAMA_URL,
        json={
            "model": "qwen2.5:7b",
            "stream": False,
            "format": "json",
            "messages": [
                {
                    "role": "system",
                    "content": (
                        "You manage order status. Reply ONLY with JSON: "
                        '{"new_status": "<status>", "reason": "<short reason>"}. '
                        "Allowed statuses: awaiting_stock, partially_cancelled, "
                        "refund_pending, shipped."
                    ),
                },
                {"role": "user", "content": json.dumps({"id": order_id, "status": status, "items": items})},
            ],
        },
        timeout=120,
    )
    resp.raise_for_status()
    return json.loads(resp.json()["message"]["content"])
Enter fullscreen mode Exit fullscreen mode

Plug ollama_decide into agent_update_order from Pattern 2 and you have a real, slow, non-deterministic agent writing to a real database. Which is exactly the environment where these bugs live.

“More agents” is not a strategy, and there’s data now

Part of why I got into trouble is that I added the second agent because it felt like the natural next step. Everything I was reading framed multi-agent as the upgrade path.

The research is more cautious than the hype. Google Research ran a controlled study across 180 agent configurations. On some tasks multi-agent setups helped a lot: on Finance-Agent, a centrally coordinated setup improved performance by about 80.9%. On others it was actively harmful: on PlanCraft, a sequential planning task, multi-agent variants degraded performance by roughly 39% to 70% depending on the architecture. A follow-up in Nature Machine Intelligence extended this to 260 configurations across six benchmarks and five architectures, and the headline didn’t change: whether more agents help depends heavily on the task structure and how the agents coordinate.

That matches my experience exactly. Tasks that decompose into independent pieces (parallel research, analyzing separate documents) benefit. Tasks where every step depends on shared, evolving state get worse, because you’ve added coordination overhead and new ways for writes to collide, which is the whole subject of this article. JIN’s piece on the four multi-agent architectures and how a supervisor dispatches work is useful here: the supervisor pattern is essentially a way of centralizing writes, which is why it tends to behave better on stateful tasks.

On the governance side, Meta’s “Agents Rule of Two” is worth knowing. The idea is that, until prompt injection is solved, an agent session should have at most two of these three properties: it processes untrusted input, it has access to sensitive systems or private data, and it can change state or communicate externally. It’s framed as a security rule, but I’ve found it useful for consistency too. The third property, “can change state,” is exactly where idempotency, ordering, and conflicts live. Every agent I give write access to is an agent I now have to wrap in keys, versions, and claims. Keeping that set small is the cheapest mitigation there is.

How do you test “two agents raced and one lost”?

This was the part I found hardest to think about. My instinct was to write tests against the agent’s output. But the race isn’t in the output. It’s in the timing.

The trick that made it testable for me: take the model out of the test. Replace the LLM with a deterministic stub that sleeps. The sleep widens the race window to something you can hit every time. Then fire the same trigger from two threads at exactly the same moment using a barrier, and assert on side effects, not on text.

# test_races.py
import threading
import time
import psycopg
import pytest
from agents import effect_key, run_once, agent_update_order, EffectInProgress, ConcurrencyConflict
DSN = "postgresql://agents:agents@localhost:5432/agents"

@pytest.fixture
def db():
    with psycopg.connect(DSN, autocommit=True) as conn:
        conn.execute("TRUNCATE side_effects")
        conn.execute("DELETE FROM orders WHERE id = 4471")
        conn.execute(
            "INSERT INTO orders (id, status, version, items) VALUES (4471, 'awaiting_stock', 0, '[]')"
        )
        yield conn

def test_duplicate_trigger_fires_one_side_effect(db):
    calls = []
    lock = threading.Lock()
    def fake_refund(payload, idempotency_key):
        time.sleep(0.3) # simulate a slow payment API
        with lock:
            calls.append(idempotency_key)
        return {"refund_id": "rf_test"}
    payload = {"order_id": 4471, "amount_cents": 1899}
    key = effect_key("run-1", "issue_refund", "order:4471", payload)
    barrier = threading.Barrier(2)
    errors = []
    def worker():
        with psycopg.connect(DSN, autocommit=True) as conn:
            barrier.wait()
            try:
                run_once(conn, key, "refund", payload, fake_refund)
            except EffectInProgress:
                pass # expected for the loser
            except Exception as e:
                errors.append(e)
    threads = [threading.Thread(target=worker) for _ in range(2)]
    for t in threads:
        t.start()
    for t in threads:
        t.join()
    assert not errors
    assert len(calls) == 1, f"side effect fired {len(calls)} times"
    count = db.execute("SELECT count(*) FROM side_effects").fetchone()[0]
    assert count == 1

def test_concurrent_writers_do_not_lose_updates(db):
    barrier = threading.Barrier(2)
    outcomes = []
    def make_agent(target_status):
        def decide(order):
            barrier.wait() # both agents read the same version, then...
            time.sleep(0.2) # ...think slowly, like a real LLM call
            return {"new_status": target_status}
        return decide
    def worker(status):
        with psycopg.connect(DSN, autocommit=True) as conn:
            try:
                agent_update_order(conn, 4471, make_agent(status), max_attempts=1)
                outcomes.append(("won", status))
            except ConcurrencyConflict:
                outcomes.append(("lost", status))
    t1 = threading.Thread(target=worker, args=("partially_cancelled",))
    t2 = threading.Thread(target=worker, args=("refund_pending",))
    t1.start(); t2.start(); t1.join(); t2.join()
    winners = [o for o in outcomes if o[0] == "won"]
    losers = [o for o in outcomes if o[0] == "lost"]
    assert len(winners) == 1 and len(losers) == 1
    status, version = db.execute("SELECT status, version FROM orders WHERE id = 4471").fetchone()
    assert version == 1
    assert status == winners[0][1] # the stored value is the winner's, and the loser KNOWS it lost
Enter fullscreen mode Exit fullscreen mode

The second test is the one I care about most. It doesn’t assert that a specific agent wins. It asserts that exactly one wins, the version moved exactly once, and the loser got an explicit conflict instead of silently overwriting. Before the version column, this test failed every single run. That’s the point: a race you can reproduce on demand is a race you can actually fix.

Once the stubbed tests pass, I run a smaller number of the same scenarios with the real Ollama model plugged in, mostly to catch cases where the model’s re-decision after a conflict does something weird. But the stubbed version is the one in CI.

The checklist I use now, before adding a second agent

I keep this in the README of anything I build with agents. It’s short on purpose.

+----+--------------------------------------------------------------+---------+
| # | Question | Pass? |
+----+--------------------------------------------------------------+---------+
| 1 | Does every side-effecting tool go through an idempotency | |
| | key derived from the intent (not generated by the model)? | |
+----+--------------------------------------------------------------+---------+
| 2 | Is that key also passed to the downstream provider? | |
+----+--------------------------------------------------------------+---------+
| 3 | Is there a reconciliation job for effects stuck in pending? | |
+----+--------------------------------------------------------------+---------+
| 4 | Does every write to shared state use a version check | |
| | (CAS / optimistic locking) instead of a blind UPDATE? | |
+----+--------------------------------------------------------------+---------+
| 5 | On conflict, does the agent re-read and re-decide, with a | |
| | bounded number of attempts? | |
+----+--------------------------------------------------------------+---------+
| 6 | Can I see who asserted what, based on which version, when | |
| | two agents disagree? | |
+----+--------------------------------------------------------------+---------+
| 7 | Do I have a test that fires the same trigger twice and | |
| | asserts exactly one side effect? | |
+----+--------------------------------------------------------------+---------+
| 8 | Does the task actually decompose into independent pieces, | |
| | or does every step share evolving state? | |
+----+--------------------------------------------------------------+---------+
| 9 | Does this new agent really need write access (Rule of Two)? | |
+----+--------------------------------------------------------------+---------+
Enter fullscreen mode Exit fullscreen mode

Here’s the uncomfortable part I had to admit to myself: items 1 through 7 apply to a single agent too. One agent with a retry policy can double-refund on its own. One agent running in two overlapping workflow runs can lose its own updates. My double refund on order 4471 wasn’t even caused by the second agent. It was a single agent retrying. The second agent just produced the lost update at the same time, which made the whole thing look like “multi-agent chaos.”

That’s the real reason for the checklist. A second agent doesn’t create new consistency problems. It makes the ones you haven’t solved visible faster. If you haven’t solved idempotency, ordering, and conflict resolution for one agent, adding another is not scaling. It’s turning up the frequency on a bug you already have.

I still use multi-agent setups. I just stopped blaming the model for bugs that belong to the database, and the database, it turns out, has had answers for these for decades.

References

Tags: AI Agents, Multi-Agent Systems, Distributed Systems, Software Engineering, Artificial Intelligence, PostgreSQL, Database, LLM, Python, DotNet

Top comments (0)