Skip to content

anthropic

Anthropic adds real-time cybersecurity filtering to Opus and Sonnet

Claude News

Anthropic is deploying a new protection layer across Claude Opus and Sonnet. It automatically detects and blocks prompts that violate the Usage Policy regarding high-risk cybersecurity activities.

The filters target two categories of activity, including direct violations like ransomware development and mass data scraping. Some tasks fall under dual-use, where legitimate security research overlaps with potentially harmful actions.

For these cases, Anthropic has launched a free Cyber Verification Program. Users submit requests via the Cyber Use Case Form, with a two-business-day turnaround for approval. The program covers Claude.ai, Claude Code, the API, and Microsoft Foundry, but is unavailable on Amazon Bedrock, Google Vertex AI, or accounts with Zero Data Retention.

Related stories

  1. Oxide is pointing Claude Mythos 5 at its own firmware
  2. Claude Code's new security plugin scans your code before you commit
  3. Anthropic will bill again for requests its safeguards block
  4. Anthropic's 225 bug finds, one attack in the wild
  5. Fable 5.1 refuses the knife but heats a gas can anyway
  6. Claude Fable knocked 20 bits off most popular hashes

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.