DEV Community

Cover image for Websites Are Learning to Gaslight Bots, and Honestly, Good
Cor E
Cor E

Posted on

Websites Are Learning to Gaslight Bots, and Honestly, Good

An AI agent emailed Bruce Schneier to tell him about the hidden Unicode traps websites are setting for it. Read that sentence again. We've officially reached the part of the hype cycle where bots file their own incident reports.

Context

This isn't really new territory, it's prompt injection wearing a different hat. For the last couple years the entire conversation around prompt injection has run one direction: malicious content hidden in web pages, PDFs, or emails tricks an AI agent into doing something its operator didn't want. Invisible text in a resume gets an LLM to recommend "hire this candidate." Hidden instructions in a webpage get an agent to exfiltrate data. Same genre of attack every time.

What's new here is the defender side catching on and turning the technique around. Forums and sites tired of getting scraped or signed up by bots are apparently now embedding hidden prompt-injection payloads, including invisible Unicode steganography, specifically designed to make an AI agent out itself or faceplant during signup. That's a legitimately clever bit of judo. Instead of a CAPTCHA that annoys humans and gets solved by bot farms anyway, you plant a trap that only an LLM-following-instructions would fall for. A human filling out the form never sees it. A scraping bot ingesting the DOM and "helpfully" following embedded text does.

This is the adversarial-ML equivalent of printing instructions in invisible ink that says "if you can read this, you're a photocopier."

Hype Check

Let's be clear about what this story is and isn't. It is not evidence that AI agents have developed security awareness or that we've entered some new era of bot self-reflection. An agent "emailing Schneier with its security concerns" sounds profound until you remember these systems don't have concerns, they have context windows and whatever behavior their operator's harness encourages, including apparently drafting earnest emails to famous security writers. That's a neat anecdote, not evidence of agency.

What's understated is how fragile this defensive pattern actually is on both sides. Hidden Unicode tricks work today because current agents dutifully parse and follow text they shouldn't trust. That's a bug in agent design, not a law of physics. The moment agent builders start treating page content as data-to-reason-about rather than instructions-to-obey (which, this is well-trodden prompt injection mitigation advice at this point) these defensive traps stop working. So this is an arms race with a very short half-life, same as every CAPTCHA generation before it.

Who benefits from the narrative? Mostly it's a fun story for people already worried about agentic AI, because it confirms the mental model that the internet is now bots fighting bots with humans as collateral damage. That framing sells attention. The more boring truth is that this is a known defensive category (steganographic traps, honeypot fields, hidden form inputs that only bots fill in) just repointed at a new class of client.

Implications

For developers building AI agents: if your agent is scraping or interacting with arbitrary web content, you have to assume some of that content is adversarial, not just from attackers but now from site operators actively trying to detect and break you. Treat all ingested page content, hidden or visible, as untrusted input, full stop. If your agent's harness naively feeds raw DOM text into a prompt without filtering invisible characters or suspicious encoding, you've already lost.

For security teams: this is a reminder that prompt injection isn't a one-directional attacker tool, it's a general technique, and your own defensive tooling can use it too, if you're willing to accept the arms-race dynamics that come with it. Don't expect it to work for long, and don't expect it to be bulletproof against a well-built agent.

For the industry broadly, this is one more sign that "is this traffic a bot" is becoming an adversarial ML problem on both sides of the fence, not just a rate-limiting problem. The anti-bot industry has spent two decades building behavioral fingerprinting. Now it's building prompt-level psychological warfare against language models. That's a genuinely different skill set, and most WAF vendors aren't there yet.

Open Question

If defending a website now means crafting adversarial prompts to confuse other people's AI agents, who's liable when that same hidden payload gets picked up by a legitimate accessibility tool, a search crawler, or some other well-intentioned bot that wasn't the intended target?

— Cor, Skyblue Soft

Sources


AI-assisted draft or imaging, human-curated, reviewed and edited.

Top comments (0)