At least $1B for outside evaluators inside Anthropic

Anthropic wants people who do not work for Anthropic grading Anthropic's models, and it has put a number on that: at least $1 billion. The company said on X that it is partnering with Accenture on independent evaluation of frontier AI, and that both companies expect to invest at least that much over the next five years.
At a glance
- Accenture is the partner named for independent evaluation of frontier AI, work that sits under Anthropic's recent commitment to embed evaluators at Anthropic, meaning the graders sit inside the company building the models.
- Both Anthropic and Accenture expect to invest at least $1 billion to build capacity in this area over the next five years, a floor rather than a disclosed budget line.
- According to unite.ai, the work covers evaluating and red-teaming models, alignment assessments and testing model safeguards; the announcement does not say how the money splits or when evaluators start.
If you have not been following, the partnership lands on top of a commitment Anthropic made recently to embed evaluators at Anthropic. Independent evaluation in the ordinary sense means an outside party runs its own tests on a model and reaches its own conclusion, rather than reading the lab's write-up. Embedding changes the geography of that arrangement: the outside party works from inside the building.
The announcement itself is short. A partnership with Accenture on independent evaluation of frontier AI, at least $1 billion expected from both sides over five years, and the stated purpose given as building capacity in this area. The post on X points to a longer item on Anthropic's own site.
The shape of the work comes from elsewhere. According to unite.ai, it will include evaluating and red-teaming models, conducting alignment assessments and testing model safeguards. Those are three separate jobs, and they sit at different distances from the model.
Red-teaming is the adversarial one: people deliberately push a model toward behaviour it is supposed to refuse, to find where the guardrails give way. Think of the locksmith you hire to break into your own building, where the useful output is a list of doors that opened. Alignment assessments ask the broader question of whether a model's behaviour matches what its builders intended, and safeguard testing checks the protections wrapped around the model rather than the model itself.
What the announcement does not do is show the seams. It does not say how the at least $1 billion divides between the two companies, how many evaluators arrive, or when the first of them start. In our view the word «independent» is carrying real weight here, because evaluators funded by both parties and embedded in one of them are a different arrangement from an outside auditor, and nothing in the post sets out who decides what gets published.
When the first evaluators arrive
No start date and no headcount have been given. The only clock in the announcement is the five-year window attached to the money, which is long enough that the interesting question is what surfaces in between. Watch for whether findings from this work ever appear in public form, and in whose name they appear, because the commitment as stated covers capacity and scope, not disclosure.
Related stories
- A UK front company sourced 1,000+ Claude accounts a month
- Anthropic says Moonshot's Kimi K3 was distilled from Claude
- Anthropic's Akamai bill could reach around $20 billion
- Anthropic will bill again for requests its safeguards block
- Trump adviser's memo puts Amodei at the root of EA
- Amazon's new plugin runs a seller's store from Claude
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
