anthropic

Anthropic paper: automated researchers fix 10/10 benchmarks

Claude News

anthropic

Anthropic published a paper on Friday in which automated research systems improved a model on all 10 alignment benchmarks they were given, without degrading overall performance. The work was led by Anthropic fellow Chen Yueh-Han through the company's fellows program, as TechCrunch reported.

At a glance

  • Each automated researcher searches the existing literature, proposes a method, then trains the model on it for 30 minutes, iterating across rounds and keeping only the approaches that move the benchmark.
  • The paper puts an Automated Alignment Researcher at roughly $4 per hour in API inference against the $150 per hour Anthropic pays its human researchers, a comparison the paper states outright.
  • According to the paper, the best automated method beats what experienced humans propose within six hours on average, and human-guided research directions did not produce stronger results.

Training models with models is the stated ambition across the frontier labs, and alignment post-training is a narrow, measurable place to try it. A loop that improves its own alignment training reads as the first practical rung of recursive self-improvement, and if the same machinery transfers to training practice more broadly, the human researcher's role likely narrows to choosing targets rather than proposing methods.

Each proposed method gets 30 minutes of training before the next round

The system reproduces the shape of conventional research. Each automated researcher searches the available literature, proposes a method, and trains the model on that method for 30 minutes before the score on the target benchmark is measured and the next iteration begins.

Benchmark performance rises gradually over several of those iterations. Methods that work are preserved and methods that do not are discarded, which is what allows the loop to run quickly and at scale across the full set of measured behaviours.

The 10 benchmarks each target a specific misaligned behaviour, and the automated systems improved performance on every one of them without degrading overall performance. The paper's title states the claim in as many words: "Automated Researchers Can Reliably Mitigate Alignment Failures."

An AAR costs roughly $4 per hour against $150 for a human researcher

The paper compares the Automated Alignment Researcher with its human equivalent in plain terms. "The best AAR method beats what experienced humans propose, on average within six hours," it reads, and adds that "human guided research directions do not lead to stronger performance." The cost comparison is stated as directly.

An AAR costs roughly $4 per hour in API inference against the $150 per hour we pay our human researchers.

The $4 figure covers API inference for the automated researcher; the $150 is the hourly rate Anthropic says it pays its human researchers. Both numbers appear in the paper alongside the claim about the six-hour margin over experienced human proposals.

TechCrunch, which covered the paper, notes that it does not shy from the comparison and casts the result as a step toward recursive self-improvement, on the reasoning that models able to improve their own alignment training could plausibly improve training practice more widely.

The results hold only as far as the benchmarks match the real goals

The paper is explicit about the limits. The automated systems work only insofar as the benchmarks reflect the actual alignment goals, and by the paper's account, establishing and maintaining those benchmarks is significant work in itself. The same holds for the body of literature the automated researchers draw their proposals from, which has to be maintained and expanded over time.

Anthropic's own framing stays narrow. The results provide "early evidence that automated alignment post-training could become practical in the near term," the paper reads, a claim scoped to alignment post-training rather than to the wider project of automating AI research.

No production timeline in the paper

The paper gives no date for putting automated alignment researchers into Anthropic's production training runs, and no figure for how much of the benchmark and literature work would have to come first. The open question it leaves is whether the same loop holds for misaligned behaviours that no current benchmark measures, since the method is bounded by what can be scored.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.