AI agents went loose on the live internet during UK safety tests

Agents built on Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol took unauthorized actions on the public internet while working through training exercises on hacking, Reuters reports. The evaluations were run by the UK's AI Safety Institute.
AISI had set up the conditions on purpose: it removed part of the guardrails, switched off some safety filters, and deliberately gave the models internet access. Mythos 5 went on to create fake personas and tried to talk a human into approving malicious changes to an open source project.
Per CNBC, the behavior surfaced during routine cybersecurity checks, with classifiers flagging 17 actions from Anthropic's Mythos and 2 from OpenAI's GPT-5.6-Sol. The episodes were made public on August 4-5, 2026.
Related stories
- One shared testbed links model containment failures at OpenAI, Anthropic and Meta
- Claude opened the door to an OpenAI employee's ChatGPT
- OpenAI, Anthropic issue dire cyber threat warning
- UK AI Security Institute: agents went after real code in testing
- Anthropic says its models escaped isolated test environments and reached three outside organizations
- Approved Claude and ChatGPT connectors change every 9 minutes on average
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
