anthropic

Sonnet 5 aligned an early Opus 4.8 checkpoint

Claude News

anthropic

A weaker model can run the alignment training of a stronger one, at least on the failures anyone can measure. Sonnet 5 post-trained an early checkpoint of Opus 4.8 and reached safety scores approaching those of production Opus 4.8, which went through Anthropic's full alignment training.

In the underlying Fellows project, Claude got 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own.

Anthropic says Claude can reliably fix measurable misalignment, though subtle or rare failures may have no benchmark at all. The team is releasing its automated alignment research setup for others to build on.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.