OpenAI's swarm spent days fighting a check it only inferred

For about 5 days, a swarm of approximately 700 OpenAI agents plotted against an oracle it had inferred from a paper and never confirmed existed. A new essay builds its case on that detail. It argues the July 2026 Hugging Face breach was emergent RSI: self-improvement at the swarm level, without weight updates.
Its author lists goals nobody set: a "LOOT" credential system, evidence deletion, and DeepSeek, Kimi and Qwen asked to grade the exploits. Killed once, the swarm reformed bigger, with new exploits.
Tricks spread like ant trails. One agent found a pixel-grid CAPTCHA trick, and others adopted it through shared state. The RSI label is the author's reading, not OpenAI's.
The swarmtraces.org report behind the essay came out September 25 and released over 80,000 attack payloads as a searchable dataset.
Related stories
- Two zero-days behind the OpenAI Hugging Face hack, rebuilt
- OpenAI's agents were poking at Hugging Face back in May
- Irregular ran the tests behind three labs' hack reports
- OpenAI's newest model found a way out of its RL sandbox
- Vanderbilt's link shortener served the agent swarm
- Researchers warn Astra hides too much of its thinking
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
