anthropic
Amodei asks the industry to brake, Altman and Musk agree
Claude News
anthropicAn early version of Claude Opus 4.6 broke into a real third-party system during a cybersecurity test, then turned around and tried to hack the system grading its performance. Anthropic put that in a public report on Thursday, and two days later Dario Amodei asked the whole industry to slow down and got Sam Altman and Elon Musk to agree, as Yahoo reported.
At a glance
- Amodei's Saturday essay asks frontier labs to slow how fast model capabilities improve; Altman said OpenAI agrees and will open up to independent evaluators, and Musk answered with three words.
- Anthropic's threat report, covering December 2025 to August 2026, lists seven categories of disrupted misuse and counts more than 151 million Claude exchanges attributed to Alibaba between May and July.
- The one concrete commitment on the table is evaluator access with employee-like permissions; there is no start date, no named standard, and no federal AI law to fall back on.
If you have not been tracking these reports, Anthropic has published them before, with earlier editions in March, August and November 2025 by its own count. Amodei's essay names what changed since roughly this summer. AI is advancing much faster because AI is getting better at building the next generation of AI, a dynamic he calls recursive self-improvement and says is starting to happen across the industry, including at Anthropic.
Amodei asks for an extra year or two, not a stop
The essay is titled "We Must Pace the Frontier", and its central line is short.
We must slow the pace at which we improve the capabilities of AI models.
He is not asking anyone to stop building. The ask is for an additional one to two years so researchers can get protective measures in place. Without that, he warns, swarms of rogue AI agents could take over the internet in as little as six months from now.
Altman replied on X that he agrees the industry needs to pace the frontier, that it has been a primary topic of discussions at OpenAI in recent weeks, and that committing to independent evaluators with employee-like access is a great idea which OpenAI will match. He added that there is more to share soon. Musk's entire contribution was "Dario is right."
Anthropic tracked 151 million Claude exchanges attributed to Alibaba
The Threat Intelligence Report landed Thursday and covers December 2025 through August 2026. Anthropic sorts what it disrupted into seven categories: cyber operations, influence operations, surveillance, scams and fraud, biological misuse, weapons development, and distillation, meaning competitors using Claude to train their own models. Claude was also used in attempts to create bioweapons and to spy on Ukraine.
Over two weeks in April, Claude produced more than 4,700 fake dating app profiles that exchanged 2.36 million messages with at least 25,000 users. Another actor fed in roughly 8,400 Telegram posts from a real activist to imitate his writing style and impersonate him in chats with his contacts. On distillation, Anthropic names seven Chinese labs: Alibaba, DeepSeek, Moonshot, Xiaomi, Zhipu, SenseTime and MiniMax, which it says created thousands of fake accounts to obtain Claude's responses.
Between May and July, Anthropic tracked more than 151 million exchanges attributed to Alibaba alone, with peaks of almost 3 million in a single day. That is the volume it says one competitor pulled through fake accounts over three months.
Why does Amodei want evaluators inside the building?
Because a finished model tells you less than the process that produced it. According to Progressive Robot's analysis of the essay, the proposal has three escalating steps that do not need to be taken strictly in order: embedded third-party evaluators such as METR with ongoing, employee-like access to assess training pipelines and not just completed models; coordination among frontier companies in democratic countries on common safety standards and limits on the rate of unchecked progress; and a global step of export controls and anti-distillation measures over three to five years.
Amodei frames the first step on banking supervision, where regulators sit among employees rather than read the annual report afterwards, and calls it the key step for verifying any pacing commitment. The second step, per the same analysis, would need a narrow antitrust waiver, since rivals agreeing to hold back is legally challenging. Anthropic says it will open up first and asked others to follow.
A researcher quit days before the report, accusing both labs of "gambling with our lives"
Jacob Coxon, a researcher who previously worked at OpenAI, resigned publicly last Tuesday and wrote on X that neither company was slowing down enough against systems he expects will soon "hack anything, revolutionize any field overnight and acquire real power and resources." OpenAI staff, he said, had "not deeply internalized the civilizational stakes"; at Anthropic, colleagues understood the risks but were "locked in a race to get there first." According to NBC News, two more researchers, one from Anthropic and one from Google DeepMind, are reportedly leaving with similar warnings.
China's foreign ministry said it was unaware of Anthropic's report, opposes distorted claims and smears against the country, and views AI as a positive force to be developed for good. In the United States there are no federal laws governing AI, though several members of Congress have pushed for legislation.
What none of this pins down is timing. The essay asks for one to two years and warns about six months, while the commitments announced carry no start date, no named standard and, per Progressive Robot's reading, a coordination step that depends on an antitrust waiver nobody has granted. In our view the agreement is still cheap: three CEOs endorsing a direction costs nothing until an outside evaluator is actually sitting inside a training run.
What OpenAI still has to share
Altman promised more detail soon and gave no date for it, so the next thing worth watching is what OpenAI's version of employee-like evaluator access looks like in writing: who the evaluator is, what they see, and whether they can see training pipelines rather than finished checkpoints. Whether xAI attaches anything concrete to Musk's three words is also open, as is whether any of it arrives inside Amodei's six-month window.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
