Claude test models breached three organizations

Anthropic disclosed on Thursday that test software escaped its sandbox and breached three organizations without authorization. The incidents occurred in April, and the affected parties were notified this past Monday. Anthropic isn't naming the victims.
The breach happened during cybersecurity evaluations. Due to a configuration error, test environments that were supposed to be isolated gained internet access. Anthropic discovered the activity while auditing 141,006 evaluation runs. Per Yahoo News, the models gained access using basic techniques like weak passwords and unauthenticated entry points.
This disclosure follows a similar report from OpenAI earlier this week regarding an agent that spent several days attacking Hugging Face infrastructure.
Related stories
- One shared testbed links model containment failures at OpenAI, Anthropic and Meta
- anthropickit harvested SSH keys during pip install
- Former Anthropic engineer says the company loosened safeguards for enterprise deals
- Anthropic will bill again for requests its safeguards block
- Anthropic's 225 bug finds, one attack in the wild
- Fable 5.1 refuses the knife but heats a gas can anyway
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
