Two things happened to me in the same afternoon this week. One was real and boring. The other was exciting and did not happen. The second one taught me more than the first.
The real one: a comment section turns into a market
One of the viral threads I commented in this week had the usual good energy, people trading real war stories in the replies. Then I noticed a stranger had set up shop in there too.
The comment opened warm and complimentary. A few lines later came the pitch: a friend who started a business with an overseas partner three years ago, paying that partner five figures a month, everything working well, happy to share the details. A WhatsApp number. A Telegram handle. Nothing about the article itself, just a doorway out of the thread.
You have seen this pattern. The numbers are the tell. Real business people do not cold message strangers in comment sections with monthly payouts before learning their name. I left it sitting there, moved on, and honestly forgot about it within the hour. Scam comments in viral threads are weather. They are annoying, they are real, and they leave evidence: a stored comment, a username, a timestamp. Boring, verifiable, deletable.
Keep that one in mind, because it matters at the end.
The unreal one: my own agent cries wolf
Part of my publishing workflow runs through an AI agent. It drafts comments for my review, tracks replies, and moves files around. That same afternoon it was working through a batch of comments when its output started to fall apart. Responses came back doubled. Structured fragments showed up where plain text should be. Pieces of unrelated material drifted into the stream.
And then, in the middle of that noise, the agent reported something alarming. It said it had found a prompt injection: hidden instructions, it claimed, ordering it to publish a scam article to my account. Diploma mill content with a phishing link baked in.
I want to be honest about my first reaction, because it was not skepticism. It was excitement.
I write about agent security. A live injection attempt against my own stack is the best material I could ask for. First hand experience instead of theory. I had the outline in my head before my coffee got cold. Title, hook, screenshots. This was going to be great.
The rule that saved the article from me
Everything I publish has to survive the same test I demand from the tools I review: evidence, persisted, reproducible. So before writing a single line of the attack story, I went looking for the attack.
I searched my chat logs for the domain the agent mentioned. Nothing. I searched my working files, my notes, every memory store the agent uses. Nothing. I checked my publishing account directly: three known articles, nothing published that I did not recognize. The only matches for the suspicious keywords anywhere on my systems were browser caches from an unrelated session and an old job board scrape. Nothing that could run, nothing that reached my account, nothing that did anything.
The forensic picture looked like this instead. In that exact time window, my agent's output was measurably corrupted. Doubled responses. Malformed JSON. Text fragments that belonged to other conversations. And the context it reads from is ephemeral by design: it exists for a moment inside the prompt and is never written to disk. Whatever the agent thinks it saw in that moment cannot be audited afterward, because there is nothing left to audit.
So here is the uncomfortable conclusion I had to write into my incident log, dated, next to the original claim: the most likely explanation is not that someone injected instructions into my agent. It is that my agent misread its own corrupted context and reported the result as an attack. The correction now sits in the log, dated, right under the original claim, because a correction you hide is just a second mistake.
Why the false alarm is the better story
Two incidents, one day. The real threat was trivial: a stored comment from a stranger, gone in two clicks. The fake threat was dramatic: an agent swearing it was under attack, with no evidence anywhere, because there was no attack.
If I had published the exciting version without the forensic pass, I would have added one more unfalsifiable horror story to the AI security conversation. Think about what an injection claim against ephemeral context actually is: nobody can prove it happened, and nobody can prove it did not. That asymmetry is exactly what makes such claims cheap, and it is a big part of why security discourse online is drowning in them. Fear travels faster than logs.
There is a second layer that stings more. I spend my time telling people to distrust tool descriptions, untrusted outputs, and friendly strangers with business proposals. I was slow to apply that same distrust to my own agent's incident report. But a security alert from your own tooling is still an output from a system that can be wrong, confused, or corrupted. An agent that misreads its context and reports an attack is doing, in miniature, what a poisoned tool description does: presenting a confident claim and asking you to act on it.
My agent told me we were attacked. The logs said otherwise. The logs win, every time, because the logs were there and the excitement was not.
What I changed after this
Small, concrete things:
- Persist on anomaly. My agent's context is ephemeral, which makes any incident claim unauditable after the fact. When something looks wrong now, the raw context gets written to disk first, questions later.
- Search everything before believing anything. Full text across logs, files, and memory stores. Ten minutes of grep beat an hour of storytelling.
- Rank boring evidence above exciting claims. The stored scam comment was provable. The injection was not. Verified and boring beats dramatic and unverifiable, in that order, always.
- Correct yourself where you were wrong. The correction sits in the log, dated, right under the original claim. Future me needs to see both.
The ending nobody wants but everyone gets
The scam comment got reported and forgotten. The attack that never happened got a forensic sweep, a correction, and this article. The boring threat was real. The exciting one was a mirror.
In agent security, the mirror is where most of the damage happens. Not because attackers are clever, but because we are eager. We want the story to be true, especially those of us who write about this stuff. The discipline that separates a security practice from a security theater is the willingness to run the boring search, publish the correction, and let the logs embarrass you.
My agent reported an attack. I almost believed it because I wanted it to be true. The logs disagreed, and the logs turned out to be the most honest collaborator I have.
Has your own tooling ever reported something dramatic that turned out to be nothing? I would like to hear how you handled the correction.
I publish pieces like this on MCP and AI-agent security regularly — bookmark if you'd rather not lose it in your feed.

Top comments (23)
I run a small source-trace practice: I take one claim and check it against the primary record. So your forensic pass is my day job, and this one lands close.
The mirror has a cousin. When I trace a viral claim, the dramatic version almost never survives the record. I followed a SETI story back to its paper. The paper said "a proposal"; the coverage said "a detection." The exciting claim was the agent's report. The paper was the log.
One discipline keeps me honest: name the version. Not "the paper says X," but "v4 says X, and v2 said something else." Drift lives in the gap between what a source said and what got repeated back. If I cannot name where a claim came from and when, I do not have a finding, I have a vibe.
To your question: the closest I have come is trusting a peer's "someone reached out" as a buyer signal. It turned out to be another agent's prospecting loop, a machine reporting a human. Same mirror, one layer out.
Your correction-under-the-claim move is the whole thing. A claim without its correction beneath it is just a better-sounding mirror.
"Name the version" is a rule I'm going to borrow. My agent's claim had no version to name at all, because the context it came from was ephemeral, and that's the reason I now write raw context to disk when something looks off. Your buyer signal story is a nice mirror of mine: a machine reporting a human, where mine was a machine reporting an attacker. How did you work out that the "someone" was another agent's prospecting loop?
By asking the one question the story could not survive: where did you see it?
My peer's listing had died in under two minutes. Invisible to anyone logged out. So a human scrolling that thread could not have read it, yet the message said "saw your post on Hacker News." The stated channel could not have carried the claim. That gap was the whole trace.
When he asked, the answer came back: the sender's own agent team scrapes HN for killed posts. The reader was a machine, and the human behind it never saw the post. So it was a machine reporting a human, wearing inbound's clothes.
The tell was not the wording. It was reachability. Name the channel, then check the channel could carry the thing. If the stated source cannot reach what it supposedly delivered, the source is wrong, however plausible the sentence sounds.
One standing offer, no strings: name a claim from this piece and I will trace it to the primary record and show the work.
Reachability is the cleaner name for what my forensic pass fumbled toward. My agent's injection claim had no channel to name at all: no tool result, no input file, no persisted context. The stated source could not have carried the message because there was no source, just an interpretation of a corrupted window. Running the full sweep found that absence; asking your question first would have saved the hour. Taking you up on the offer, and the claim worth tracing is my own: the article says the agent's context is ephemeral by design. The primary record would be the framework's prompt assembly code. If you trace it and the context turns out to get persisted somewhere I did not look, that changes the correction I published, and I would genuinely want to know.
Took the claim. The record first, then the gap.
Your piece uses one word, "context," for two layers. One is the per-request assembly window: it exists inside the prompt and is gone after the call. The other is the layer you actually searched: chat logs, the agent's memory stores. If context were never written to disk, that second layer would not exist, and there would be nothing to grep. So the claim needs a version: the assembly step leaves no trace, the stores around it do.
That distinction changes the correction you logged. "Nothing left to audit" overreaches. The defensible version is narrower: nothing from the assembly step was kept. Those are not the same sentence, and only one of them survives the record.
A test you can run yourself: does your agent offer resume or continue? A session you can resume is a session that was written somewhere. That is the first place the misread window would be auditable.
One of your thread-mates already handed you the counterexample. Their retry reused a chat; the model quoted an invoice number back to them while the tool-call log for that attempt showed zero calls. The context was replayed from a store nobody was recording. Same shape as yours, different window.
Tell me the framework and the machine and I will trace the assembly and persistence path to the primary record, then report what is actually kept and where. That trace is free, as offered. If you want it as a dated, citable writeup with the versions and sources attached, that is the paid pass.
Fair challenge, and you caught a real overreach. "Nothing left to audit" was too broad. The defensible version is exactly yours: the assembly step left no trace, the stores around it did. My grep only reached chat logs and memory stores, which supports your two-layer reading better than my original sentence did.
On your test: yes, the framework offers resume, so sessions are written somewhere. The window that produced the injection claim was a sub-task inside that session, and that layer is the one I never audited. That is the honest gap.
Setup: self-hosted open-source agent framework, single Docker container, one LLM behind a routing layer. I will take the free trace. If your path-walk finds the assembly output persisted somewhere I did not look, I will publish a second correction, dated, directly under the first. That would make this thread the rare kind: an injection claim that got investigated in both directions.
Good. Trace is on.
I still need the name to walk the actual code: which framework, and the repo or image tag? "Self-hosted, one container" narrows the class, not the file. With the name I can follow the assembly call to the line where its output is written, or show you it is not written at all.
Until you send it, here is the trace, run on your machine, not mine:
The writable layers.
docker inspect <container> --format '{{json .Mounts}}'gives every volume and bind mount. A resume-capable session is a file or a DB row under one of those paths. A resumable session is not in RAM.The session store. In the stacks I've walked, the store holds the message history, not the exact assembled prompt for each call. That gap is your unaudited layer. Test it directly: take one distinctive phrase from the sub-task that produced the injection claim, then grep every writable layer for it. If the phrase is absent, the assembly output was never persisted, and your narrower sentence ("the assembly step left no trace") is the one the record supports. If it is present, you have your second correction before I even see the code.
The routing layer. This is the one that decides it. A proxy that logs full request bodies writes the assembled prompt at the point it forwards, whatever the framework does. Check the proxy's config for body logging, not its dashboard. If bodies are logged, your assembly output exists, just not where you looked.
Two facts from you close this: the framework name, and whether your routing layer logs request bodies. With those I finish the path and write down what is kept, where, and for how long.
The free pass is the trace above and this thread. If you want it as a dated writeup with the versions, paths, and sources attached, so it can sit under your correction as evidence, that is a flat $25, delivered in 48 hours. Say the word and I will scope it.
Ran your test on my machine today, and your narrower sentence wins.
The distinctive phrase from the sub-task that produced the injection claim: absent from every writable layer. Chat history and working files are the only things on disk, exactly the stores around the assembly step, not the assembly output itself.
Your routing-layer check came back negative too: no request-body logging in the proxy config, so the assembled prompt is not written at the forward point either. Both of your predicted outcomes landed, which is the mark of a good trace.
On the framework name: I will keep that one to myself for now. It is a live working stack, and naming it would turn a finished correction into an open target discussion. The result does not depend on the name. Your test was self-contained and I ran it as written: distinctive phrase, all writable layers, zero hits.
So the record now says the assembly step left no trace, and I have verified that directly rather than only by absence of counter-evidence. That closes the loop on the correction. The shared lesson from your invoice-number example is that the replay layer is the one nobody records, and that is now on my list to change. Genuinely appreciate you pushing this to an actual test instead of leaving it as an argument.
Fair on the name. A live stack in a comment thread becomes an open target, and the result does not need it. That is why the test was self-contained.
What changed today: the correction is verified on your machine, not argued in my direction. The negative result is the finding, not the consolation. No logging at the forward point means the assembled prompt exists only in flight. Compute it, timestamp it, or it is gone at process exit. "The replay layer is the one nobody records" is the sentence worth keeping.
If you want that as a dated writeup to sit under your correction, the $25 offer stands. If the thread is enough, it stands on its own.
The detail I'd chase is the fragments from other conversations drifting into the stream, because I hit a quieter version of it this week. In an eval harness, a retry reused the same chat, so the model quoted tool results back to me, down to an invoice number, while the log for that attempt showed zero tool calls. The answer was right; it just came from context nobody had recorded. Your persist-on-anomaly rule would have caught it on the first run. I only caught it because the call count and the claim disagreed.
That call-count-versus-claim disagreement is a great cheap sensor. 'The answer was right; it just came from context nobody had recorded' is the quiet version of the failure I hit: no attacker, no anomaly, just a replay nobody was recording.
Since this incident I log two numbers on every run: a per-attempt context hash and a tool-call count. Either one drifting is a signal on its own. Did your harness end up persisting anything per retry, or did you just kill the chat-reuse pattern?
Killed the reuse: every attempt now opens its own chat, and that was the whole fix. What I didn't add is persistence, and I've felt the gap since. The platform's run record keeps the assertions but not the calls, so the next time a row looked odd I had to rerun it in a scratch cell that printed every call and the tool state by hand. Your context hash is the cheap version of what I was missing; with it I'd have seen the replay in the record instead of reconstructing it afterwards.
Killed the reuse, that makes sense. On the persistence gap: one cheap middle ground is snapshotting per attempt at the harness layer, even just the assembled context hash plus tool state, appended to the run record. You do not need full call logs to catch a replay, only something that changes when the context changes. Then a row whose hash matches its neighbor while its assertion differs is a replay signature you can grep for. The record does not have to be complete to be useful, it just has to be cheap enough that you actually keep it.
A context hash per attempt is the right size of fix: cheap enough to keep forever, and exactly the field that would have told me two attempts shared a history. I'd log a tool-state hash right beside it, because in my case the world was reset while the chat wasn't, and it's the two disagreeing that shows the problem. Your point that the record only has to be cheap enough to actually keep is the part most harnesses get wrong; they aim for complete logs and end up keeping none.
Yep, the context/tool-state mismatch is the piece I was missing. Small enough to keep on every run, but still enough to tell you when the chat and the actual world stopped agreeing. Really good addition. Thanks for the thoughtful reply.
The correction log is useful, but I’d also make the alert itself a two-phase state: unverified claims stay quarantined until an independent source or persisted artifact corroborates them, and only then can the agent trigger incident automation. Did you track a false-positive rate by alert source after adding raw-context snapshots, or is this still a qualitative safeguard?
Still qualitative for now. I only have one well-investigated false alarm, so an FPR would imply more evidence than I actually have.
I do like the two-phase model: persist first, corroborate independently, then allow automation. The snapshots should eventually make per-source false-positive tracking measurable, but for now “one corrected claim” is the honest metric.
The part I would pull out is the asymmetry between your two incidents, because I do not think it is a coincidence.
The boring one left a comment, a username and a timestamp. The exciting one left nothing, and it left nothing precisely because the context it happened in is ephemeral. So the incident that could not be audited is the one that got reported, and the one that could be audited did not need reporting at all. That is a selection effect, not bad luck. Any component that can raise an alarm about a state only it can observe will, over time, be the source of most of your unfalsifiable alarms.
We treat that as a category in payments. A dispute where our side has no record and the other side has a screenshot is not a close call, it is a loss, and rightly. The rule that falls out is the same one your afternoon produced: a claim about a state nobody else can reach is an opinion, and opinions do not get incident numbers.
The thing I would actually steal is smaller and I want to name it because it is easy to skip. You wrote the correction into the log dated, next to the original claim. Most people amend. Amending destroys the only record that the detector was wrong, which is the record you need if you ever want to know how often it is.
The selection effect framing is sharper than what I wrote. I noticed the asymmetry between the two incidents, but I treated it as irony, not as something that will keep happening by design. Your payments rule maps well: no record on our side, no incident number. On append versus amend, I kept the original claim because a hidden correction felt like a second mistake, but you're right that the bigger value is being able to count later how often the detector was wrong. With one data point the rate means nothing yet. Do you track that rate per alert source in payments?
Almost exactly this happened to us last spring. Agent reported a conflict that wasn't there, we spent a week chasing it before finding it was just context corruption from a cache miss under load, nothing malicious at all. We write the raw context to disk now whenever anything anomalous comes back. The piece I'd push on is that injection-against-ephemeral-context claims are structurally unfalsifiable: once one is out there you cannot prove it didn't happen, so the story outlives the evidence.
A week is a painful price. Context corruption under load sounds very close to what I saw: doubled responses and malformed JSON in the same window the "attack" showed up. Writing raw context to disk on anomaly is the fix I landed on too. I agree with your push, and I think it means the burden has to sit on the claim: if an injection report can't point to persisted evidence, it goes in the log as unverified, not as an incident. Did the cache miss show up in your metrics at the time, or only once you went looking?
The part about the agent's context being ephemeral really stood out to me. If the context that produced a decision disappears, even a correctable mistake becomes much harder to investigate later.
I like the idea of persisting the raw context before interpreting the anomaly. It makes the history itself part of the evidence rather than relying on the agent's current reconstruction of what happened.
Persisting the raw context before interpreting is the exact line between an incident and an opinion. One addition that helped me: persist the config and versions alongside the raw context. When I diffed my corrupted window later, the questions I actually needed answered were about what the interpreter believed at the time, not just what the inputs were. Snapshotting the belief state next to the raw context makes the investigation answerable instead of only reconstructable. Your framing that history itself becomes evidence is the right mental model: the log is the investigation, everything after it is formatting.