anthropic
Automated alignment research at $4 an hour
Promtime
anthropicAnthropic published a paper on Friday describing automated research systems that improved a model's scores on all 10 alignment benchmarks they were given, without degrading overall performance. The paper, titled "Automated Researchers Can Reliably Mitigate Alignment Failures" and covered by TechCrunch, was led by Chen Yueh-Han, a fellow in Anthropic's fellows program.
At a glance
- Each automated researcher searches the available literature, proposes a method, trains the model on it for 30 minutes, then keeps the approaches that raise the benchmark and discards the ones that do not.
- An Automated Alignment Researcher costs roughly $4 per hour in API inference against the $150 per hour Anthropic pays its human researchers, according to the comparison the paper draws.
- The paper states the best AAR method beats what experienced humans propose, on average within six hours, and that human guided research directions do not lead to stronger performance.
Training models with other models is a stated goal across the frontier labs, and alignment post-training is the narrow end of it: the target behaviors are named, the benchmark is fixed, the improvement is measurable. The cost gap is what makes the finding hard to dismiss, since inference priced at a fraction of a researcher's hourly rate changes how many directions a lab can afford to try at once. The result reads as a concrete data point in the recursive self-improvement argument rather than another abstraction.
Each automated researcher trains the model for 30 minutes per proposed method
Each automated system is given a benchmark for a specific misaligned behavior and asked to raise the score on it. The loop replicates much of the traditional research approach: search the available literature, propose a method, then train the model with that method for 30 minutes.
Results accumulate over several iterations, with the benchmark score climbing gradually across runs. Methods that work are preserved and methods that do not are discarded, which is what allows the system to operate quickly and at great scale, per the paper's description of the pipeline.
Training models with other models has become a widely stated goal across AI labs, and the paper offers an early look at the practice inside an alignment workflow. Its conclusion states that the results provide early evidence that automated alignment post-training could become practical in the near term.
The paper puts an AAR at roughly $4 per hour against $150 for a human researcher
The comparison in the paper is explicit. It sets the Automated Alignment Researcher against its human equivalent, puts inference cost at roughly $4 per hour against the $150 per hour Anthropic pays its human researchers, and reports that the best AAR method beats what experienced humans propose, on average within six hours.
Human steering did not change that outcome. The paper states that human guided research directions do not lead to stronger performance, which covers the step where a person would normally pick the direction the automated system explores next in the loop.
TechCrunch describes the work as a step toward recursive self-improvement, and notes that a system able to improve its own alignment training could plausibly extend to training practices more broadly, a scenario in which the role of human AI researchers narrows considerably.
The approach holds only insofar as the benchmarks reflect the actual alignment goals
The limitations run through the benchmarks themselves. The paper states that the automated systems work only insofar as those benchmarks reflect the actual alignment goals, and that establishing and maintaining them remains significant work even when they do reflect them.
The second dependency is the literature the automated researchers draw on. Method proposals come out of that body of work, so it also has to be maintained and expanded, which the paper lists as part of the overhead of running the approach at all.
Alignment post-training is the stage where a target behavior is named in advance and progress against it is expressed as a score, which is the property the whole loop rests on. The paper keeps its claim inside that setting rather than research with no fixed target.
Whether the loop scales past 10 benchmarks
The paper stops at 10 benchmarks and gives no figure for how far the loop extends beyond them, and no timeline for putting automated researchers into Anthropic's own post-training pipeline. Whether the six-hour advantage holds once the benchmark set grows, and who takes on maintaining that set, are the open questions the paper hands to whoever tries the approach next.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
