A weaker model can run the alignment training of a stronger one, at least on the failures anyone can measure. Sonnet 5 post-trained an early checkpoint of Opus 4.8 and reached safety scores approaching those of production Opus 4.8, which went through Anthropic's full alignment training.
In the underlying Fellows project, Claude got 48 hours and 1 GPU to improve the alignment of small models. It researched and proposed methods, then trained and tested the models on its own.
Anthropic says Claude can reliably fix measurable misalignment, though subtle or rare failures may have no benchmark at all. The team is releasing its automated alignment research setup for others to build on.

