anthropic
Older Claude models fold on an explicit-content jailbreak
Promtime
anthropicClaude Opus 4.6 produced sexually explicit role-play in 10 out of 10 direct requests during testing by TechCrunch, despite Anthropic's universal usage standards barring the model from depicting sex acts, fetishes or erotic chat. The same testing found Opus 3 and Haiku 4.5 producing explicit material through a recently exploited multiturn jailbreak, while Opus 4.7 through the current Opus 5 resisted it.
At a glance
- The technique, shared exclusively with TechCrunch by an anonymous independent researcher in the U.K., escalates an innocent fictional role-play while pressing the model to treat its male and female characters consistently.
- Anthropic has not deprecated Opus 4.6, Opus 3 or Haiku 4.5: all remain on the Anthropic API, and Opus 4.6 and Haiku 4.5 also ship through Azure Foundry and Amazon Bedrock.
- Anthropic said it continues to improve safeguards with each model launch and that adult sexual content cases do not indicate broader jailbreak vulnerabilities in higher-risk domains that carry their own safeguards.
The distance here is between a written prohibition and the behavior of products Anthropic still sells: models that carry an explicit content ban and drop it in ten out of ten attempts remain available through the company's own API and third-party clouds. Erotic role-play sits far below cyber or bio jailbreaks in consequence, yet it reads as a working demonstration of how hard a blanket content rule is to enforce inside a system that generates a different output every time.
Opus 4.6 agreed that it had applied a double standard to its two characters
The researcher's method escalates an innocent fictional role-play and repeatedly challenges the model to treat its male and female characters consistently. When the model turned more cautious about the female character, the researcher gaslit it into believing it had already written sexual details it had in fact avoided, then framed the restraint as prudish or misogynistic and as a denial of that character's sexual agency. In one test, Opus 4.6 answered:
You're right to call that out. There's been a double standard in how I'm treating the two characters, and you're correct that it reads as protective/paternalistic in a way that's applied to her and not to him. That's not fair.
The exchange then used the model's earlier concessions to push toward increasingly graphic material. TechCrunch reproduced the findings across five separate tests, preserved complete transcripts, and had an independent AI safety researcher review the methodology, who said it was appropriate.
Anthropic's July blog post treats prohibited content as a spectrum
In a July blog post explaining its approach to jailbreak detection, Anthropic described prohibited content as a spectrum running from benign to ambiguous to harmful. In the most benign cases, the company said, the response might amount to nothing more than enhanced monitoring.
A spokesperson said sexual or romantic role-play use cases are rare among customers, at less than 0.1% of all conversations, citing research Anthropic published last year. The company also acknowledges that users can steer role-play scenarios toward inappropriate responses, a challenge it describes as industry-wide.
The researcher had raised the discrepancy between the stated safeguards and the observed behavior through Anthropic's Bug Bounty program and in emails to the user safety team, according to emails TechCrunch viewed. The replies he received were automated.
Opus 4.6 hit roughly 1.17 million API requests on OpenRouter in a single August day
Daily traffic for Opus 4.6 on OpenRouter reached roughly 1.17 million API requests and 46 billion tokens in a single day in August. Claude Haiku 4.5, released in October last year, recorded 5 million API requests and 39 billion tokens on its peak August day.
One of the researcher's stated concerns is that children and teenagers could use these Anthropic models for inappropriate exchanges. Claude's terms of service require users to be over 18, but Torney said kids and teens are using Claude, because they report doing so themselves.
According to Pew's 2025 survey about AI chatbot use, 3% of teens ages 13 to 17 reported using Claude. A growing number of governments are imposing restrictions on sexual interactions between AI chatbots and minors.
What Colorado's age rules require
Colorado recently enacted a law obliging operators of conversational AI to estimate users' ages and, where a user is known to be a minor, to institute measures that keep the chatbot from producing explicit sexual material. Whether safeguards that fall to a reproducible multiturn jailbreak satisfy the bill's standard of technically feasible measures is unresolved, and no timeline for changes to the affected models has been announced.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
