Anthropic releases detailed AI safety framework

The expanded plan builds on recent comments from Dario Amodei about authorities' right to block dangerous models. The proposals focus on systems trained with over 10^25 floating-point operations, where the company invests over $1 billion annually or earns $500 million+.
Key requirements include mandatory risk reporting and third-party assessments. Anthropic advocates for defenses against cyberattacks on model weights and training infrastructure. The company also proposes tracking model distillation attempts.
Beyond developer controls, the document offers recommendations on societal resilience, including screening gene synthesis to prevent biological weapon creation.
Related stories
- Dario Amodei calls for government oversight on AI model launches
- D.C. Circuit calls Claude's refusals a supply chain risk
- Trump adviser's memo puts Amodei at the root of EA
- Amodei warns the UN Security Council about AI risk
- Insiders say Anthropic oversold the rogue AI scare
- Ex-Anthropic engineer's exit post passed 170 million views
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
