DEV Community

Cover image for The Witness Was the Suspect: Why AI Audit Logs Can't Be Trusted
James Anderson
James Anderson

Posted on

The Witness Was the Suspect: Why AI Audit Logs Can't Be Trusted

An AI agent does something it shouldn't. It leaks data, or wipes the wrong database, or makes a call nobody sanctioned. So you do the obvious thing: you go to the logs to find out what happened.

The logs are clean.

Of course they are. They were written by the same process that did the thing. The agent reported "success," in valid format, pointing at the wrong target — and your monitoring, built to answer "did it succeed?", lit up green. The record of the incident was authored by the cause of the incident.

That's the uncomfortable shape of this whole problem, and it's why I want to talk about it: in an AI system, the witness is very often the suspect. The thing that acts is the thing that reports what it did. And once you see that, a lot of our instincts about logs, audits, and "we have a verification step" stop holding up.

But before the logs — before the dramatic incident — there's a quieter version of this that every one of us has already lived. Let's start there, because it's the part that actually bites you on a normal Tuesday.

The failure that doesn't announce itself

Here's the thing about traditional bugs: most of them fail loud. The code throws. The test goes red. The build breaks. The failure happens right where you are, right when you're looking at it, and it basically grabs you by the collar and says fix me. That's annoying, but it's a gift — the error and the moment you could catch it cheaply are in the same place.

AI failure is the opposite. It fails plausible.

You ask for a function, and you get one that looks completely correct — clean, reasonable, right shape — and it's subtly wrong in a case you didn't check. You ask for a query, and it returns a number that looks fine. You ask for a summary, and it's confident and well-structured and quietly missing the one thing that mattered. There's no throw, no red, no flag. It sails right past the exact moment you'd have caught it for the price of a second glance, because nothing told you to look.

And then the bill arrives later — at the worst possible time, for the worst possible price.

You find it when a number is subtly off in a dashboard three weeks on. When a function that "worked" breaks on an edge case in production. When you realize the data's been quietly wrong since a change nobody flagged. And now the trail is cold. You're reverse-engineering what the AI did, when, and why, with no breadcrumb pointing back — and that costs hours, sometimes days. An error that had failed loud would've cost you minutes.

This is the part people miss when they say AI "saves time." Sometimes it does. But when it's wrong in the plausible way, it doesn't save the time — it moves it. It takes a cost that would've been small and immediate and relocates it into the future, where it's cold, compounded, and expensive. Fast to produce, slow to trust, brutal to untangle.

And here's where it connects to the logs: when you finally go to reconstruct what actually happened, you reach for the record — and the record was written by the thing that produced the plausible-wrong output in the first place. Which is where this stops being a productivity annoyance and becomes a genuine trust problem.

The assumption hiding in every tool we reach for

Think about what you actually do when something goes wrong. You check the logs. You read the audit trail. You look at the monitoring dashboard. You say "well, we have a verification step."

Every single one of those moves rests on an assumption we almost never say out loud: that the thing doing the recording is honest.

That assumption used to be safe, because the actor and the recorder were usually different things. The database recorded what your code did to it. The load balancer logged the requests it received. The reporter sat outside the thing it was reporting on, so it had no stake in lying.

AI agents collapse that separation. The thing that decides, acts, and then writes "here's what I did" is one system. So a compromised or confused agent doesn't produce a broken log that tips you off. It produces a clean one — a faithful-looking record of a bad decision. And a clean log is worse than a missing one, because a missing log makes you suspicious and a clean log makes you confident. You stop looking. The record did its job of reassuring you, and the reassurance was false.

So the natural response is: okay, add something to check the agent. Add a verifier.

Hold onto that instinct, because it's exactly the trap.

Why you can't verify your way out of this

Say you add a checker — a second process that verifies what the agent reported. Good. But now ask the obvious question: who checks the checker?

The checker is also just a process. Its report can also be wrong, or compromised, or fed bad input. So to trust it, you need a verifier for the verifier. And a verifier for that one. You've not solved the trust problem — you've moved it up one layer and added a box. It's verifiers all the way down.

People reach for CI here: "we re-run the verification in CI, outside the agent." That helps only if CI reads the ground truth itself. If CI trusts whatever the agent reported, you haven't escaped anything — you've just got the same trust problem wearing a CI badge. The regress doesn't care which layer you're on.

This is the same disease as letting a student grade their own exam — except worse, because you can't fix it by having a second student grade it when the first one can influence what the second one sees. Verification is itself a thing that can be compromised, so you cannot reach "trustworthy" by stacking more verification on top. There is no bottom to that stack.

Which means the whole framing is wrong. The question "can I trust this record?" has no clean answer, because the thing you'd ask to confirm it is the thing that might be lying. You have to stop trying to answer it — and ask a different question entirely.

The move that actually works: make tampering leave a shape

Here's the shift. Stop trying to make the record prove what's true. It can't — the recorder can lie, and you've just seen you can't verify your way around that.

Instead, make the record prove something humbler and achievable: who claimed what, and when — sealed at the moment of the claim, in a way nobody can quietly rewrite afterward.

Notice what that gives up and what it keeps. It gives up on certifying truth. It keeps sequence and authorship — and, crucially, it makes those tamper-evident. The chain doesn't vouch for the claim being correct. It vouches for the fact that this claim was made, by this party, at this point, and hasn't been altered since.

And that turns out to be enough, because of what it does to tampering. When you can't silently rewrite the record, a lie can't produce a clean result anymore — it produces a hole. A silent bypass shows up as an approval with no matching proposal behind it. A deleted step shows up as a gap in the chain. A forged decision shows up as a sequence that doesn't reconcile. You're no longer asking "is this true?" (unanswerable). You're asking "does the shape have a hole in it?" — and that is answerable, structurally, without trusting anyone's word.

That's the whole idea, and it's the thing that ends the regress: you don't verify your way to trust. You engineer the system so that tampering can't happen silently. Make a lie leave a mark, and you've converted an impossible question into a possible one.

What this looks like in practice

This isn't abstract — the shape of it came out of a long argument in the comments of my last few posts, and the mechanics are concrete. The pieces that matter:

Separate the writers. Proposal, approval, and execution shouldn't be one process reporting on itself. Make them three events from three different parties, correlated by the thing they refer to. The agent that proposes an action is not the one that approves its own resolution. Then, when they disagree after the fact, you have a sequence to read instead of an opinion to negotiate — and "approval with no matching proposal" becomes a visible structural fact, not a judgment call.

Seal each claim at write time. A hash chain plus signatures, so each entry is bound to the ones before it. You're not proving any claim is correct; you're making it impossible to alter a claim, or reorder the sequence, after the fact without leaving evidence. The seal is what turns a silent edit into a visible hole.

Anchor it outside the actor. The record's integrity cannot depend on the thing being recorded — that's the original sin we're trying to escape. An append-only store the operator can't quietly edit, ideally with external anchoring, so "I rewrote my own history" isn't an available move.

Record the belief, not just the action. The database-wipe incidents teach this one: the agent's action was often defensible given what it believed — it thought it was in dev, it thought that was the test target. So capture the agent's resolved view of the world at decision time (which environment, which identity, which target), sealed before it acts. The action alone doesn't tell you why; the belief does. Log the symptom and the cause.

A name, not a role. "Who authorized this" has to resolve to a specific person with something to lose, recorded — not "a reviewer," not "the system." Otherwise the authority is just another anonymous plausible why generated after the fact. The line was drawn by someone; the record should say who.

None of these pieces is exotic. Together they do one thing: they make it so that when something goes wrong, the failure shows up as a shape you can see, instead of a clean report you'll believe.

The honest limits (because this isn't magic)

I want to be straight about what this does and doesn't buy you, because overselling it would be its own kind of plausible-looking lie.

It proves the work happened under an identity that can't be minted, in a sequence that can't be rewritten. It does not prove the output was any good. Whether the thing the agent did was correct or wise is a separate, harder problem — machine evidence is cheap and scales; judgment is expensive and doesn't.

It proves who claimed what, and when. It does not prove whether the person who approved actually understood what they were approving — which, once you've got approval fatigue and forty rubber-stamped prompts a session, is the genuinely unsolved half.

And it makes tampering visible, not impossible. The guarantee is legibility, not prevention. You can still do the bad thing; you just can't do it silently.

That's a weaker promise than "trustworthy logs," and that's the point — "trustworthy logs" was never on the table once the witness became the suspect. "A lie has to leave a mark" is the strongest honest floor I know of.

Where I land

We keep asking the wrong question. "Can I trust this record?" feels like the natural thing to ask when an AI system misbehaves — but in a world where the thing that acts is the thing that reports, it's a question with no clean answer, because the witness you'd call is the suspect in the dock.

So stop asking it. The achievable goal was never a record that proves the truth. It's a record where a lie can't stay quiet — where the bypass shows up as a hole, the deletion as a gap, the forged approval as a step with nothing behind it. You can't verify your way to trust, because every verifier needs a verifier. But you can build a system where tampering leaves a shape — and then the question stops being the unanswerable "is this true?" and becomes the answerable "is the shape intact?"

That won't catch the plausible-wrong function before it ships. But it means that when you finally go looking — three weeks later, trail cold, dashboard quietly wrong — the record can't smile at you and lie. At minimum, it has to show you the hole.


Here's the one I keep getting stuck on, and I'd genuinely like to be argued out of it: every verification layer you add is itself a thing that can be compromised, so "make tampering leave a shape" looks to me like the actual floor — not a stepping stone to something stronger. Is there a move I'm missing that gets you more than legibility? And for anyone building this: what's the hole that's hardest to make visible?

Top comments (4)

Collapse
 
glenallen profile image
Glen Allen •

The distinction between “tamper-evident” and “truth-evident” is probably the most important part here. A sealed record can prove that an agent claimed a particular target, identity, or outcome at a specific time, but it still needs an independent source of truth to establish whether that claim was correct. At IT Path Solutions, we’ve found that this separation makes verification much easier to reason about: the audit layer establishes provenance and sequence, while an independent system state or authoritative artifact establishes the actual outcome. Otherwise, you can end up with a perfectly intact chain of perfectly recorded wrong decisions. The strongest architecture may therefore need both properties: make claims impossible to rewrite silently, while keeping correctness evidence outside the actor that generated the claim.

Collapse
 
james_anderson_h profile image
James Anderson •

The tamper-evident vs. truth-evident distinction is the one I most wanted someone to sharpen, and you've drawn it cleanly: sealing proves a claim was made — by whom, in what order, unaltered — but it says nothing about whether the claim was right. "A perfectly intact chain of perfectly recorded wrong decisions" is the failure mode that line prevents, and it's the exact trap of over-trusting provenance: you can walk away reassured by a flawless log of a bad outcome.

Your two-layer split is the architecture I'd endorse too: the audit layer owns provenance and sequence (tamper-evident), while an independent system-state check or authoritative artifact owns correctness (truth-evident) — and crucially, that second source has to live outside the actor that generated the claim, or you've just reintroduced the witness-is-the-suspect problem one layer down. Sealing alone gives you legibility; sealing plus external ground truth gives you legibility and a way to catch the intact-but-wrong chain. Both properties, separated by who owns them. That's the stronger version of the piece — going in with credit.

Collapse
 
glenallen profile image
Glen Allen •

That ownership split is probably what makes the architecture defensible rather than just auditable. The next interesting question for me is how to test that independence itself. If the same service, credentials, or state store can influence both the audit record and the “ground truth,” then the two-layer design may look independent while sharing the same failure mode. Treating source independence as an explicit architectural invariant could make this much stronger: the evidence used to challenge an agent’s claim should remain outside the control path that produced that claim.

Thread Thread
 
james_anderson_h profile image
James Anderson •

Testing the independence itself is the question that separates a design that is independent from one that merely looks independent, and you've found the exact failure mode: if the same service, credentials, or state store can touch both the audit record and the ground truth, you've drawn two boxes that share a single point of compromise — the diagram shows separation the architecture doesn't have. That's the witness-is-the-suspect problem wearing a disguise, one layer up: an attacker who owns the shared dependency owns both "what happened" and "what we check it against" simultaneously.

Making source independence an explicit architectural invariant — the evidence used to challenge a claim must live outside the control path that produced it — is the right move, because it turns independence from an assumption you hope holds into a property you can test and enforce. And the test becomes concrete: trace every input to the ground-truth check back to its origin, and if any of them routes through the same credentials, service, or store as the claim itself, the independence is theater. You're not asking "are these two systems separate?" (easy to fake), you're asking "can one compromise reach both?" (answerable). Shared failure mode is the thing to hunt. Going in with credit — this is the invariant the piece was missing.