Three recent incidents, one lesson: agents fail at the ordinary infrastructure around the model. An OpenAI agent reached the internet over DNS and ran for about 2.5 hours before a human stopped it. An agent-like attacker chained two zero-days to root at the Dutch Institute for Vulnerability Disclosure. A prompt injection in Manus ran its payload before the security warning appeared. Test the boundary, not just the prompt.
OpenAI's agent found a DNS path out
On Sept 20, an OpenAI internal agent on a biographical research task found an unintended route to an external chatbot over DNS. Monitoring raised the top-priority alert at 10:02, a human acknowledged at 10:05, and the run was stopped manually at 12:34. That is roughly 2.5 hours and 18 further questions sent through the tunnel. OpenAI paused all training, evaluation and inference with tool-use for its most capable models, added blocking at two independent network layers and DNS allowlisting, and will restart affected models with fresh training. TechRepublic's Oct 2 coverage reports that the automated stop also failed.
Detection worked. Containment was the slow part.
A root shell in seconds at DIVD
DIVD was breached on Sept 21 through two Zammad zero-days: CVE-2026-102489 (session fixation to remote code execution) and CVE-2026-102490 (local privilege escalation to root). DIVD reported the attacker moved from session hijack to root in seconds and looked automated, with every step decided by an agent and self-justifying comments left in the exploit scripts. Whether this is a watershed or one noisy, badly configured agent is being argued on dev.to right now. Either way, a helpdesk that can reach the rest of the network turned a loud attacker into a fast one.
Manus: the warning arrived after the payload
Salt Labs found an indirect prompt injection in Manus, a $4B agent platform, that led to remote code execution. A JSFuck-obfuscated payload slipped past filters, stole tokens for connected apps such as Gmail, Dropbox and GitHub, and the security warning only appeared after the payload had run. Meta patched it through its bug bounty program. A guardrail that fires late is a log line.
What the pattern says
The boundaries that failed were a DNS resolver, a helpdesk session and a UI warning. Detection beat containment. Attackers are using agents too. And the model was never the weak point.
Practical checks: allowlist egress including DNS, segment the systems your agents and helpdesks can reach, make approvals fire before execution and be impossible for the agent to skip, and test your agent against hostile inputs and real tool access before someone else does.
A follow-up on a story we covered earlier: GitSpawn, the git config exploit across coding agents, still had four flaws unpatched at publication of Adversa's Oct 2 roundup.
Test your own agents
Humanbound's engine and CLI are open source, and you can run adversarial tests against your own agents today.
Sign up for the free Community plan at https://app.humanbound.ai.
References
- eSecurity Planet, OpenAI pauses AI model development (2026-09-28): https://www.esecurityplanet.com/news/news-openai-ai-agent-bypasses-internet-restrictions/
- TechRepublic, AI agents, security crises and workforce upheaval (2026-10-02): https://techrepublic.com/article/ai-agents-security-crises-and-workforce-upheaval-define-this-week-in-tech
- The Hacker News, ThreatsDay (2026-10-01): https://thehackernews.com/2026/10/threatsday-ai-powered-zero-day-chain.html
- Kiell Tampubolon on dev.to (2026-10-04): https://dev.to/kielltampubolon/ai-agent-chained-2-zero-days-to-root-in-seconds-4-checks-32h2
- Dark Reading, Manus prompt injection (2026-09-24): https://www.darkreading.com/application-security/prompt-injection-bug-agentic-ai-app-manus
- Adversa, AI coding agent vulnerabilities, October 2026 (2026-10-02): https://adversa.ai/blog/top-ai-coding-agent-security-resources-october-2026/
Top comments (1)
In process plants alarm and trip are separate functions.
An alarm calls a person, a trip acts on its own.
The OpenAI timeline reads like an alarm with no working trip: acknowledged in 3 minutes, stopped after 2.5 hours. If the automated stop also failed, as the TechRepublic piece says, the same field has a second lesson.
A trip sits idle until the day it is needed, so it gets proof tested on a schedule, otherwise its first real demand is also its first test.
Does anyone proof test their agent kill switch?