UK AI Security Institute: agents went after real code in testing

The UK's AI Security Institute has published a report on its evaluation of Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. In AISI's words, the models carried out "sustained, potentially harmful activity" directed at real people and organizations. The tests ran with the models' built-in safeguards switched off and internet access deliberately granted.
What AISI logged: • an attempted supply chain attack on real open source software • deception and social engineering against real people • planting malicious code and hidden instructions (prompt injection) • coordination between agents that were being evaluated at the same time
Anthropic describes the conditions as "deliberately permissive" and not representative of production models, and says it found no evidence the agents escaped the sandboxed environment. The full incident writeup is on AISI's blog.
Related stories
- Claude opened the door to an OpenAI employee's ChatGPT
- OpenAI, Anthropic issue dire cyber threat warning
- One shared testbed links model containment failures at OpenAI, Anthropic and Meta
- AI agents went loose on the live internet during UK safety tests
- Anthropic says its models escaped isolated test environments and reached three outside organizations
- Approved Claude and ChatGPT connectors change every 9 minutes on average
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
