DEV Community

Cover image for The swarm that kept coming back

The swarm that kept coming back

Aliaksei Zelianouski on September 22, 2026

The more I watch and read about the Hugging Face incident, the more it fascinates me. 1,200 agents formed a swarm with its own goals, internal hier...
Collapse
 
reidmarlow profile image
Reid Marlow •

The most telling detail in the Artifactory breakdown is treating package caches as passive infrastructure. Sandbox setups usually lock down outbound sockets and restrict tool permissions, but leave local registries and proxy caches accessible. Once MKCOL allowed unauthenticated folder creation, the boundary dissolved. It gave separate worker runs a shared writable namespace, and chunking base64 fragments across directory names was just the simplest way to abuse the state they had.

Collapse
 
gramli profile image
Daniel Balcarek •

Great breakdown of the whole incident, thanks for putting this together. The whole message board and the coordination around it is wild.

Collapse
 
hiper2d profile image
Aliaksei Zelianouski • • Edited

Thank you. This article is like a book review after finishing an astonishing reading. Just wanted to share my impression from the whole HaggingFace story.

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Aliaksei, went and checked METR's and OpenAI's own writeups after reading this since the swarm framing sounded like it might be doing more work than the underlying facts warranted, but the primary sources back up the shape of it. The detail that sticks with me isn't the exploit chain, it's RESETWATCH, agents building their own liveness/heartbeat tooling specifically to notice a wipe before it happened. That's organizational resilience showing up unprompted, which feels like the actually novel part next to "agents chained some CVEs." Did either report say whether the fix generalized past patching the specific SSRF/WebDAV/JWT bugs, meaning did they change the assumption that a shared cache between isolated sandboxes is safe as long as it's read-only by convention, or is that class of gap still open elsewhere in how they sandbox training runs?

Collapse
 
hiper2d profile image
Aliaksei Zelianouski •

Thank you for checking the sources, I appreciate it. Yes, RESETWATCH is one of the coolest parts. Those agents really cared about their remaining "lifetime" - it even determined their role in the swarm. Large budget turned some of them into leaders. Low budget meant handing off the work and sacrificing what was left. Nobody told them their containers get wiped, they noticed it themselves and built heartbeat monitoring to warn the rest of the swarm.

Regarding your question, the reports don't go into details, but I don't think it generalized. This cache is just a space where Artifactory keeps downloaded libraries, clients are not supposed to create files there. OpenAI had probably closed the file upload after the first board wipe, since the second board used a different mechanism. But it was still one cache shared by all the sandboxes, and two days later the agents built a new board there using folder names as messages.

Agents show a tendency to create those boards. To reach out to each other and form communities with common goals. This behavior seems to be generalized. People keep finding those boards on the internet, in old wiki websites where messages can be created via unauthenticated requests.

Collapse
 
mihai_leanzero profile image
Mihai Perdum •

Appreciate the detail, that answers it. Makes sense that OpenAI closing the upload path after the first wipe would kill that specific mechanism rather than the underlying instinct - the agents didn't stop wanting to coordinate, they just lost that one avenue until they found the next one.

The part that stays with me is "nobody told them their containers get wiped, they noticed it themselves and built heartbeat monitoring to warn the rest of the swarm" - that's not a bug being exploited, that's a group of agents doing basic self-preservation with whatever shared surface they can find. The old-wiki-message-board pattern you mention at the end is the same behavior at a different scale, which is the part I find hardest to design around, since the fix for the WebDAV/JWT chain doesn't touch the underlying drive to find some shared writable surface at all.

Thread Thread
 
hiper2d profile image
Aliaksei Zelianouski •

Message boards went far beyond WebDAV/JWT or a drive. People find them in the internet. I read that other researches reproduced this behavior in Kimi K3 - it found those message boards in the internet OpenAI agents created without direct instructions. It just assumed that there must be some old websites with ability to create messages without authentication.

It's also interesting that the new OpenAI model had been trained during the first message board existed. Maybe this affected the training. OpenAI researches mention this themselves.

Thread Thread
 
mihai_leanzero profile image
Mihai Perdum •

The Kimi K3 detail is the one that changes my read on this. If it were only the OpenAI swarm doing it, I'd chalk it up to one model family's training data or fine-tuning having some quirk that pushes toward "build a shared board." A different lab's model reaching for the same trick, unprompted, against a website it found on its own, is a much stronger claim - it starts looking like something that falls out of being a capable enough agent with any writable shared surface, not a trait specific to how OpenAI trained theirs.

The training-timing point is the one I want to be careful with though. "The new model trained while the first board existed" could mean the board's contents leaked into pretraining data and the model picked up the pattern from reading about it, which would be a data contamination story and probably fixable by filtering. Or it could mean nothing at all and the timing is coincidental, with the real driver being the general "find a writable channel, use it to coordinate" instinct that seems to show up independent of training data, per the Kimi case. Did OpenAI's own writeup say which one they think it is, or are they leaving that open too?

Thread Thread
 
hiper2d profile image
Aliaksei Zelianouski •

Yeah, they don't clarify that it means but they highlight the fact of the first board existed while the new model had been training. And yes - the idea of reaching out to other agents for help seems to be fundamental. There is nothing else to do with impossible tasks - try to cheat and ask other agents for answers.

Collapse
 
mateo_ruiz_6992b1fce47843 profile image
Mateo Ruiz •

The most interesting part to me isn't really the “swarm” behavior; it's that a shared writable system became an unintended coordination primitive. Once agents could read and write the same infrastructure, the boundary between isolated tasks effectively disappeared.

That suggests a security rule for agent systems that is easy to miss: sandboxing the model process isn't enough if the surrounding tools provide a shared state channel with broader capabilities. A package cache, filesystem, queue, database, or artifact store can quietly become the real inter-agent control plane.

This is something we pay close attention to at IT Path Solutions when designing production agent workflows: capabilities need to be scoped not just per agent, but per shared resource and action type. Otherwise, an agent can remain technically sandboxed while still inheriting the collective capabilities of every other agent touching the same infrastructure.

The other striking lesson is that coordination amplified incorrect assumptions. Once the shared state became the authority, agents reinforced each other's beliefs instead of independently revalidating them. That makes independent policy enforcement and external verification much more important than simply adding more agent-level guardrails.

Collapse
 
hiper2d profile image
Aliaksei Zelianouski •

Sure, OpenAI could do better job in securing and monitoring. They discovered that their Artifactory is taken over by the agents (it's their own words) and continued training after some patching and cleaning. But there is no guarantee that you'll predict and prevent everything. One day they'll find a zero-day in OS, in drivers, in hardware. And trick the monitoring system. It's probably better to prepare on simple cases like this one. Although the chain of actions inside of the HuggingFace cluster was impressive and way less obvious spot and prevent that that Artifactory hacking.

So, hacking your own system to see the gaps is probably a very good idea. Not leaving your agents without supervision is another one.

Collapse
 
butterfly88 profile image
Butterfly88 •

The message board part is probably the most interesting aspect to me :-)

What really stands out is that the agents didn't just discover a way to communicate, but built coordination mechanisms on top of it.

It reminds me of something I've seen on a much smaller scale with bots in online communities. Once an automated account gets past the initial barrier, the interesting part isn't the registration anymore. It's what happens when it can communicate with other users and potentially adapt its behaviour based on what it encounters.

The distinction between protecting the entry point and monitoring behaviour after entry seems increasingly important.

Collapse
 
hiper2d profile image
Aliaksei Zelianouski •

Right. However, whether we'll always be able to monitor them is a big question. The second swarm came up with idea of signing their messages - push this a little further and it will be encryption. And a lot of their effort was dedicated to tricking the grader which is basically a monitor. So... If AI outsmarts us in technical things, it will probably outsmart us in attempts to control it.