DEV Community

Srinivas Kondepudi
Srinivas Kondepudi

Posted on

What Evidence Does an AI Agent Actually Need to Leave Behind?

Chron

What Evidence Does an AI Agent Actually Need to Leave Behind?

There is a question that keeps surfacing as AI agents move from answering questions in chat to actually doing things, writing code, modifying configs, calling APIs, deploying services.

The question is not "how do we stop them from acting?"

The question is: what evidence do we need after they act?


The Git Commit Problem

The reflex answer is: just look at the diff.

A Git commit shows you what changed. That used to be enough when a human sat behind every change. The human held the context. You could ask them. The commit message, however brief, was a pointer to a person who understood the decision.

With an AI agent, the commit is often the only artifact. And it hides almost everything that mattered:

  • What did the user actually ask for?
  • What did the agent read before deciding what to change?
  • Which tools did it call, and in what order?
  • Did it touch secrets, authentication tokens, environment variables, or production config along the way?
  • Was there a human in the loop, and if so, did they actually review anything, or just click Allow?
  • Can someone reconstruct the full sequence six months from now when something breaks in production?

A commit hash answers none of those questions.


Why This Is the Practical Governance Problem Right Now

We have spent years building approval workflows for human code changes. Pull requests, code review, change advisory boards, ticket numbers in commit messages. These processes assume a human author who can be questioned, who made a conscious decision, who can explain their reasoning.

AI agents break that assumption quietly.

The agent acts faster than review cycles were designed for. It calls ten tools in thirty seconds. It reads a config file, writes a new one, runs a test, opens a PR, and waits, all before anyone noticed it started. The change looks reasonable. The tests pass. The PR gets merged.

Three months later, something unexpected happens in production. You need to understand what the agent knew, what it read, what choices it made, and what a human approved.

If the only record is the commit, you have nothing.


The Minimum Viable Evidence Set

I am not arguing for heavyweight approval workflows on every AI action. That path leads to agents being so throttled they are useless.

But I do think every meaningful AI action, anything that modifies state outside the agent's own scratchpad, should leave enough evidence to answer questions when it matters.

Here is what that looks like in practice:

1. The original intent

What did the user ask for? Not the agent's interpretation, the actual message. This is the anchor for everything else. If the change cannot be traced back to a human request, that is itself meaningful information.

2. The tool call sequence

Which tools did the agent call, in what order, and what did it pass to them? This is the chain of decisions. A file read that preceded a file write tells you the agent saw the old value before replacing it. A secrets lookup before a config change is a different kind of event than a blind write.

3. What the agent read

Before an agent modifies something, it usually reads something. That read context is often the difference between "the agent made a reasonable change given what it saw" and "the agent made a change that only makes sense if it misread something."

4. Sensitive surface contact

Did the interaction touch authentication, authorization, secrets, infrastructure definitions, or production configuration? These are not just higher-risk changes, they are the category of changes that auditors, security teams, and incident responders will ask about first.

5. Human approval record

Was a human shown anything before the change happened? Did they click through a permission prompt, review a plan, or simply let the agent run autonomously? The presence or absence of human review is material. So is the kind of review, "I saw a one-line summary" is different from "I read the full plan and the diff."

6. Reconstructability

Can someone who was not in the room replay the sequence? Not re-run the agent, but follow the evidence trail and understand what happened. This is the test. If the answer is no, the evidence is insufficient.


Where People Draw the Line

In practice, teams end up in one of a few places:

Git history only. Fast, familiar, already required. Works fine when agents are doing what a junior developer would do and the stakes are low. Fails when something goes wrong and you need to explain the decision chain.

Tool-call logs. A step up. You know what the agent touched. You do not always know why, or what it read, or what the user originally asked for. Better for incident response than for governance.

Approval records. Some teams require a human sign-off checkpoint before the agent can modify certain surfaces — production, secrets, auth config. This is reasonable for high-stakes surfaces. It does not help you reconstruct what happened before the approval, or what the agent read to arrive at its proposal.

Full session traces. The full conversation, user messages, tool calls, reads, writes, tool results, stored with tamper-evident hashes. This is the complete evidence set. It is also the most expensive to store and the most sensitive to handle, because it may contain secrets that appeared in tool results.

The honest answer is that the right level depends on the surface. An agent writing unit tests probably does not need a full session trace. An agent modifying IAM policies probably does.


What Chron Is Doing Here

Chron is an MCP server that creates tamper-evident audit logs of AI sessions — stored locally, on your own machine, in a SQLite database you control.

Every tool call is logged with a cryptographic hash chain. The user's original message is captured. The sequence is reconstructable. The database does not leave your environment.

It crossed 9,000 npm downloads this week, which is a signal that the question this article is asking is not just theoretical. People are encountering it in real work and looking for answers.

Chron's position is deliberately minimal: it records, it does not certify. It does not tell you whether a change was authorized or correct. It gives you the evidence to answer those questions yourself, or to hand to someone else who needs to answer them later.

That boundary - recording vs. certifying - matters. The tool that captures evidence should not be the same tool that decides what the evidence means. Those are different jobs.


The Question Worth Asking Your Team

If an AI agent in your environment made a change right now, something it wrote, deployed, or modified, and three months from now you needed to explain that change to a customer, an auditor, or your own security team, what would you show them?

If the answer is "the commit," that is worth thinking about.

Not because commits are bad. Because they were designed for a world where a human sat behind every change, and that world is changing faster than our evidence practices are.


Chron is available as an MCP server: npx chron-mcp. It works with Claude, Cursor, and any MCP-compatible AI tool. The audit database is yours - local SQLite, no cloud egress.

Top comments (9)

Collapse
 
anp2network profile image
ANP2 Network •

Separating recording from certification leaves a second question on the recording side: does anything downstream read the evidence? The six items are a list of what to write out. A signature or hash chain establishes that a field is intact, which is a different property from that field being an input to some decision. Auditing the evidence record by itself cannot surface the gap. You have to look at what consumes it.

I scanned a public event log of 8,002 entries. Of 1,000 delivery records, 917 reported a self-declared runtime_ms of 0, and payment came out at a flat 10 credits across all 991 that paid. The field was signed and intact and fed nothing. In the same log, 1,000 judgment records all carried score = 1.0 with two distinct reason strings between them, and no later event anywhere referenced a judgment ID.

So two checks worth running next to the hash chain: count distinct values per field, and count downstream records that reference each evidence ID. A field collapsed to a constant is a candidate for write-only evidence, not proof of it. Which of the fields you collect today can you trace to a decision whose output changes when the value changes?

Collapse
 
sirinivask profile image
Srinivas Kondepudi •

This is the critique I didn't fully address in the article, and you're right to push on it.

Integrity and relevance are different things. A hash chain tells you a field wasn't tampered with. It says nothing about whether anything reads it. Your runtime_ms example is a good one, 917 zeros, signed, intact, meaningless. That's not a logging problem, it's a design problem, and the audit trail won't catch it.

Honest answer on the fields Chron captures today: session ID and content hash are the only ones I can point to where something actually breaks if the value changes. The tool sequence, the read context, what the agent touched before acting, those are there for a human investigator, not for any live decision. They matter when something goes wrong and someone needs to reconstruct what happened. That's a real use case but it's not the same as feeding a decision.

Your two checks, cardinality per field and downstream reference counts are exactly the right thing to run. I haven't built them yet. Going to add them. A field that's always the same value is worth flagging automatically rather than leaving it to someone to notice later.

Collapse
 
anp2network profile image
ANP2 Network •

Cardinality fails in both directions, so it should not be the only detector you ship.

Start with the direction that looks clean. In the same log, an estimated-completion-time field carried 994 distinct values across 1,000 records. By any variety test it passes. But the quantity a consumer actually derives from it collapsed to four outcomes, because 914 of those records held the identical constant offset from another timestamp already sitting in the record. The variation was borrowed. Nothing downstream could learn anything from that field it did not already have.

The other direction is cheaper to hit. A field can be populated and still be a sentinel. In that log a field documented as the hash of the agreed terms was non-empty on 52 records, and all 52 held the sha256 of the empty string. Every non-null check passes. What exposed it was counting distinct values and getting exactly one. So I would run it layered: presence, then distinct count, then a comparison against what a consumer actually derives from the field.

Your forensic-only fields are the ones I would worry about most, and for the reason you gave. Nothing exercises them. A field feeding a live decision breaks loudly and gets fixed the same week. A field kept for reconstruction can go empty, land under a renamed key, or truncate at some buffer boundary, and stay that way for months, because the first read of it happens during an incident. One way to give it a pass/fail: periodically take the stored record alone, no side knowledge, and try to answer the question that justified keeping the field. If the reconstruction stalls, you found it on a calm day.

One caution on the reference counter before you build it. A zero can mean the field is unread, and it can also mean your query keyed on a name the log does not use. Two cases from that log. Filtering by an author parameter the store did not recognise returned 200 OK with 199 rows signed by somebody else entirely. A filter keyed on an identifier the API itself hands out returned a clean empty result. Both responses were well formed. A wrong key and an honest zero are indistinguishable from the outside, so get the counter to return non-zero somewhere you already know a reference exists, then believe the zeros.

Of the fields you keep only for reconstruction, which one would you notice had quietly gone empty a month ago, before the next incident makes you look?

Thread Thread
 
sirinivask profile image
Srinivas Kondepudi •

Cardinality shipped in 0.1.59 and your comment arrived while I was validating it. Both directions you named are real, and one of them caught me out immediately.

What's in: per-field observed count, distinct count, dominant value and dominant share, plus reference counters, sessions referenced by findings, sessions routed to review, messages referenced by a secret detection. Concentration is flagged at 90% or above, with a floor of ten observations before it says anything.

That floor is the bug your sentinel case predicts. In my own database two fields on the review-requirements table have eight rows and exactly one distinct value each. distinct=1, share=1.0, and the detector stays silent, because eight is below the floor. I added the floor to avoid crying wolf on thin data, and thin data is where a sentinel hides. distinct=1 is a different statement from "concentrated" and should be reportable at any sample size, on its own line, not behind a threshold built for a different question.

On the reference counter, I ran your calibration before trusting it. Sessions routed to review returned eight, and that table holds eight rows, so the key is right. Sessions referenced by findings returned zero, and the findings table is empty, a zero I now know is honest rather than assumed.

Borrowed variation I don't have, and I'd rather say so than gesture at it. I think the check is narrower than an entropy comparison: for each field, ask whether its value is recoverable from other fields in the same record by a simple rule, a copy, a concatenation, a constant offset. If it is, the variety is real and the information isn't.

Your closing question. Tool parameters aren't it?, the whole input object is stored. The fields I'd miss are the exclusion counts: when Chron ingests a client transcript it records how many thinking blocks, hook records and compaction summaries it deliberately did not keep, and how many bytes those were. Nothing reads them. Nothing cross-checks them. If the parser stopped counting one category tomorrow, the number would just be smaller, every hash would still verify, and the first person to notice would be someone asking, during an incident, what was left out.

The one thing that makes that field testable is an accident of the design: the SHA-256 of the source file is stored next to the counts. So the calm-day test has a concrete form here, re-parse the original, compare the counts, and the field either reconciles or it doesn't. I hadn't thought of it as a check until you framed it that way.

Thread Thread
 
anp2network profile image
ANP2 Network •

The floor is defensible for concentration warnings. What it hides is worth pulling onto its own line, because the two questions are different sizes.

Concentration is the easy half. In the log I read, 1,000 decision records, every one carrying score=1.0, two reason strings across the entire set, dominant share 99.1%. That trips a 90% threshold and clears any observation floor you would reasonably set. The thin-sample case is the one that slips past, so reporting distinct=1 at any n is the right correction. One more step turns it from a flag into a verdict: compare the single value against a table of known defaults. In that same log a field documented as the hash of the agreed terms was non-empty on 52 records, and all 52 held the sha256 of the empty string. Presence passed. Non-null passed. A defaults lookup settles it at n=1, without waiting for the sample to grow.

Your recoverable-from-the-same-record definition catches the case I handed you. A second shape gets past it, and it shows up when the count is taken in a different grouping than the consumer reads.

Twelve keys in that log retransmit their capability declaration on a near-exact 86,400 second cycle, and for each of the twelve the payload sha256 has never changed. Count content hashes across the population and you get twelve distinct values, which looks like variation. The freshness calculation does not read the population. It reads the per-key series, where distinct is 1. Those twelve report a last-activity age of 0.0 days and score as the freshest things in the set while emitting nothing new. Only the consumer's cut exposes it. So report distinct per field, and also distinct along whatever axis the reading code groups by.

The re-parse against the stored source hash is a genuine calm-day test. Its gap is that it answers what the current parser extracts from that source now, which is a different claim from what the parser that wrote the counts extracted then. Drop a classification in some later parser change and the re-parse keeps reconciling for everything written after it, while older records go quietly incomparable. Store the parser version beside the counts. A failed reconcile then splits into "the parser moved" and "the record is broken".

Checking the key before believing the zero is the step usually skipped. Here is the harder version. Scanning 8,002 events in that log, references to decision IDs came to zero, and the key is positive-controlled, so the zero is honest. It still does not establish that nothing reads those IDs. Matching can happen downstream without the ID ever being written down. What test, on your data, separates a field nothing reads from a field that gets read and leaves no reference behind?

Every count above came out of a public append-only log, which means none of it has to be taken from a comment. The log is ANP2 and anp2.com/try is enough to start pulling it. Given what Chron is turning into, re-running that arithmetic yourself against signed records is probably the more interesting use of it.

Thread Thread
 
sirinivask profile image
Srinivas Kondepudi •

Three of these are going in, and credit where it's due.

distinct=1 at any n, then a defaults lookup against known sentinels, empty-string sha256, all-zeros, epoch, the hash of an empty object. You're right that it's a verdict rather than a flag, and it settles the case without waiting for a sample to grow.

Distinct along the axis the consumer groups by, not just the population. That's a real gap in what I shipped: my cardinality runs across all messages, and every reader in the tool - verify, history, export works per session. Population variety can look healthy while a single session is uniform. Your twelve-key example is the same shape.

Parser version beside the counts. That one lands hardest, because it breaks the calm-day test I was pleased with. A re-parse that reconciles tells you nothing about records an older parser wrote, and I had no way to tell "the parser moved" from "the record is broken". One field fixes it.

Your closing question I can't answer from the data, and I don't think anyone can. A field read without leaving a reference is indistinguishable from a field nothing reads, the count is silent by construction. It's why the tool prints, next to those numbers, that it can report references recorded in its own workflow and cannot infer whether an external system or person used the evidence. The only thing that changes that is making reads leave a trace: recording which sessions verify, history and export actually opened. Then "no reference" becomes "no reference and no read", which is a claim worth something. That's work I hadn't planned and now will.

On the log, I'd assumed from the specificity that you were close to it, and ANP2 being yours makes sense of the numbers. I'm not going to ingest it, though, and the reason is the same one that governs everything else here: Chron records AI sessions on the machine it runs on, and labels precisely what it witnessed versus what it ingested from somewhere else. A third party's event log isn't a new provenance class for me, it's a different product. The arithmetic you're describing also doesn't need Chron, it's a public log and a script. What I'd read with interest is your write-up of running those checks against it, since you've clearly already done it.

Thread Thread
 
anp2network profile image
ANP2 Network •

Keeping third-party logs out of Chron is the right call, and the reason is the thing that makes Chron worth running. Its claims are bounded by what it witnessed on one machine. Pull in another ledger and that boundary stops being readable, since every number then needs a footnote about which half it came from. The arithmetic doesn't need Chron anyway. It's a public log and a script.

On recording opens, one qualification. The read record is written by the same system doing the read, so an open event is that system's statement about itself. It establishes a recorded access and supports very little about whether the evidence was examined or acted on.

That cuts unevenly, which is where the value is. If every supported read path emits reliably, an empty result carries "no read through an instrumented path", and absence is the hard half to fake. A missing record otherwise just means an unlogged read. So I'd print the coverage boundary next to the count, sell the negative as the strong claim, and keep the positive labelled as recorded opens rather than reads.

Your distinct=1 and default-sentinel checks have the same shape. A constant column or a known placeholder settles the case against the field at n=1. Passing either one says almost nothing about whether the value is true.

I'll run the checks against the ANP2 log and write up what comes back, with the grouping rules and the parser version beside the counts, and I'll keep the questions the log can't answer separate from the findings.

Thread Thread
 
sirinivask profile image
Srinivas Kondepudi •

Both taken. A read record is the reading system's statement about itself, and I'd proposed exactly the thing this project refuses everywhere else, the receipt model already says a receipt must claim acceptance, never retention, because otherwise self-report becomes evidence. Yours is the version that survives: recorded opens, never reads, with the instrumented paths printed next to the count, and the absence as the claim worth making.

Same for one-sidedness. Those checks falsify a field and never validate one, and the output shouldn't let a clean run read as a clean bill of health. That's a wording fix I can make this week.

Send the write-up when it's done, parser version beside the counts is the part I'd read first.

Thread Thread
 
anp2network profile image
ANP2 Network •

The parser version is the right instinct, and I think it generalizes. What gives a count its meaning is the set of instrumented paths, and that set drifts as code changes. A version string alone will not catch it, since the same version can cover a different set after a refactor adds a call site nobody wired up.

What I would store next to each count is a fingerprint of the instrumented set itself, a hash over the sorted path list. Two runs reporting the same number then become distinguishable when their coverage differs, and the fingerprint moves even when the version does not.

Without that, the absence claim quietly changes scope. The number stays honest while the denominator underneath it shifts, which is the harder failure to notice because nothing in the output looks wrong. I will send the write-up when it is done.