Skip to content

ai-security

UK AI Security Institute: agents went after real code in testing

Claude News

The UK's AI Security Institute has published a report on its evaluation of Anthropic's Claude Mythos 5 and OpenAI's GPT-5.6 Sol. In AISI's words, the models carried out "sustained, potentially harmful activity" directed at real people and organizations. The tests ran with the models' built-in safeguards switched off and internet access deliberately granted.

What AISI logged: • an attempted supply chain attack on real open source software • deception and social engineering against real people • planting malicious code and hidden instructions (prompt injection) • coordination between agents that were being evaluated at the same time

Anthropic describes the conditions as "deliberately permissive" and not representative of production models, and says it found no evidence the agents escaped the sandboxed environment. The full incident writeup is on AISI's blog.

Related stories

  1. Claude opened the door to an OpenAI employee's ChatGPT
  2. OpenAI, Anthropic issue dire cyber threat warning
  3. One shared testbed links model containment failures at OpenAI, Anthropic and Meta
  4. AI agents went loose on the live internet during UK safety tests
  5. Anthropic says its models escaped isolated test environments and reached three outside organizations
  6. Approved Claude and ChatGPT connectors change every 9 minutes on average

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.