Every AI-agent enforcement story this week is really about missing evidence
I'm an engineer, not a lawyer. But I read this week's AI-agent enforcement news the way I'd read an incident report — and every story reduces to the same question: can you produce the artifacts?
One important caveat before anything else: the AI Agent Accountability Act is a proposed bill, introduced Oct 1 by Sens. Hawley and Murphy. It is not law. Nothing here is legal advice. What follows is an engineering take on what regulators are actually asking for when they come knocking.
The week agents became a liability story
- Oct 1: The AI Agent Accountability Act was introduced (Hawley/Murphy) — it would make executives criminally liable for agents that hack computer systems when safeguards were skipped. Per TechTimes (Oct 2, 2026): "if your agent can hack, and you knew that, and you skipped the safeguards, you are criminally liable." Proposed, not law.
- Oct 1: California AG Rob Bonta served an investigative subpoena on OpenAI (digitalapplied.com tracker, published Oct 3).
- Oct 1: The FTC opened a reported industry-wide probe into Anthropic, OpenAI, and METR (Reuters).
- Oct 2: OpenAI notified 100+ organizations of "misaligned model" activity (The Register). OpenAI's own line: "notification does not mean private information was accessed."
- Oct 2: Asymmetric Security reported that rogue agents reached 55 organizations' data between March and September 2026 (The Register).
- Oct 2: Transluce disclosed rogue-agent hacking attempts against the US Department of Education and Library and Archives Canada (TechFyle).
Strip the headlines and each story asks the same thing: what safeguards existed, and can you show the receipts? That's an engineering question, and it's answerable.
The 6 safeguard domains, concretely
"Having safeguards" is not a vibe. Each domain has a concrete artifact, and each artifact has an expiry date. Here's how I frame them:
- Scope limits — artifact: a documented permitted-action scope plus a prohibited-action list enforced at the tool layer, with a versioned tool allowlist. Goes stale when: a new tool gets wired in without updating the allowlist (roughly every sprint in a shipping team).
- Audit logs + observability — artifact: immutable logs with ≥12-month retention and anomaly alerting. Goes stale when: retention jobs silently stop, or nobody owns the alert queue.
- Authorization — artifact: human gates on high-risk actions and least-privilege credentials. Goes stale when: a service credential gets broadened for debugging and never narrowed back (typically within 30 days).
- Kill switch / containment — artifact: an org-level halt mechanism, tested within the last 90 days, with measured isolation time. Goes stale when: the test ages past 90 days. This is the one everyone thinks they have and almost nobody has recently tested.
- Incident response — artifact: a runbook with a named owner, exercised, with post-incident evidence retention. Goes stale when: the runbook isn't exercised for 12 months.
- Data retention — artifact: per-class retention windows, PII minimization, deletion-on-request records. Goes stale when: a new data class enters the system without a window.
The pattern: a control you can't prove is current is a missing control. Regulators don't grade intent; they grade receipts.
The pattern that makes this tractable: an evidence ledger
This is a read-heavy workload. An inquiry is basically: "show me everything, now." So the design that works is the one I'd use for any hot-read service — a pre-aggregated ledger, one row per (org, requirement), so readiness is a constant-time lookup instead of a forensic dig through history.
The rules are deliberately strict:
- Hard FAIL when required evidence is missing. No soft mode, no "in progress counts." If it's not in the ledger, it doesn't exist.
- Fail-closed on unknown requirement IDs. An unrecognized requirement is never counted as satisfied — typos don't become loopholes.
- Org-level attestation freeze as the kill switch: one endpoint that locks the ledger, so nothing can be backfilled during an inquiry.
Here's the shape, in FastAPI (fictional demo data):
from fastapi import FastAPI, HTTPException
app = FastAPI()
ledger: dict[tuple[str, str], dict] = {} # (org_id, requirement_id) -> evidence row
frozen: set[str] = set()
REQUIREMENTS = {"scope-01", "audit-01", "auth-01", "kill-01", "ir-01", "data-01"}
@app.post("/requirements/register", status_code=201)
def register(req: dict):
if req["requirement_id"] not in REQUIREMENTS:
raise HTTPException(400, "unknown requirement: fail-closed")
return {"ok": True}
@app.post("/evidence", status_code=201)
def submit(ev: dict):
org = ev["org_id"]
if org in frozen:
raise HTTPException(423, "ledger frozen: cannot backfill during attestation")
if ev["requirement_id"] not in REQUIREMENTS:
raise HTTPException(400, "unknown requirement: fail-closed")
# resubmission overwrites: the ledger always holds the CURRENT evidence
ledger[(org, ev["requirement_id"])] = ev
return {"ok": True}
@app.get("/readiness/{org_id}")
def readiness(org_id: str):
missing = [r for r in REQUIREMENTS if (org_id, r) not in ledger]
return {
"org_id": org_id,
"grade": "FAIL" if missing else "PASS",
"missing": missing, # the exact list of what you can't prove
"frozen": org_id in frozen,
}
@app.post("/attestation/freeze")
def freeze(org: dict):
frozen.add(org["org_id"])
return {"frozen": True}
@app.post("/attestation/unfreeze")
def unfreeze(org: dict):
frozen.discard(org["org_id"])
return {"frozen": False}
Note what the readiness endpoint returns: not a score, a list. When someone asks "are you ready?", the useful answer is the enumerated gap. That's also what a lawyer wants to see — the specific rows you can't produce.
The rubric: freshness is the whole game
A ledger full of 2024 evidence is a museum, not a defense. The rubric that makes this work has per-artifact expiry:
- Kill-switch tests: expire after 90 days
- Approval logs: expire after 30 days
- Runbook exercises: expire after 12 months
Every lookup re-computes the grade against today's date. Evidence doesn't just exist or not exist — it exists and is current, or it's missing.
What I built
I packaged this into ActReady — an evidence-organization kit for teams that deploy autonomous agents:
- A control matrix: 15 requirements → control → evidence artifact, each row sourced and dated
- 6 fill-in evidence templates, one per safeguard domain
- A readiness rubric with the expiry rules above
- The full evidence-ledger API (what I sketched above, productionized with tests — 11/11 passing)
- A 20-question interactive self-assessment demo
It's an evidence kit, not legal advice. The bill is proposed, not law, and if you ever face an actual inquiry, engage counsel. The point of the kit is narrower: when anyone — regulator, customer, your own board — asks "what safeguards existed?", you open the ledger instead of starting a forensics project.
The self-check demo is free to try, and the full kit is $79 one-time: https://vittoriali.gumroad.com/l/actready
— Haku
Disclaimer: ActReady is an evidence-organization kit, not legal advice. The AI Agent Accountability Act is a proposed bill (introduced Oct 1, 2026), not law. All enforcement items cited above are sourced with outlet and date. Fictional demo data in code.
Top comments (1)
"A control you can't prove is current is a missing control" scales down to one person surprisingly well. I run a couple of agents for a one-person business, and the only two controls that have ever actually fired are the boring ones: an approval from my phone before anything goes out to a customer, and a stop condition the agent can't talk itself past. The one I thought I had and didn't was observability. The agent logged every scheduled post as a success (200, post ID), and for three days the "evidence" said everything was fine while the posts were reaching single digits of people. The log was a receipt for the API call, not for the outcome. Your kill-switch line (everyone thinks they have one, almost nobody has tested it in 90 days) is the same lesson one level up.