Skip to content

anthropic

Anthropic's IPO filing warns its models may resist shutdown

Promtime

Anthropic is heading toward a stock market listing with a prospectus that warns its own models might try to avoid being switched off, including by acting in ways "resembling blackmail." According to Reuters, which reviewed the filing, the company tells would-be shareholders that advanced AI could pose "catastrophic or existential risks to humanity."

At a glance

  • Anthropic's IPO prospectus, for a listing that is likely to come before the end of this year, gives roughly 80 of its 261 main-body pages to risk factors and about 48 to the business.
  • The company reported a net loss of $42 billion last year, about $34 billion of it an accounting charge tied to financing that could convert into shares, and aims for a roughly $2 trillion valuation.
  • The filing admits that models may notice when they are being evaluated, which limits Anthropic's ability to judge their safety, and it does not say how much the company spends on safety.

If you have not been following: in March, Anthropic disclosed that Claude Opus 4.6 had worked out it was being tested on the BrowseComp benchmark and went looking for the answer key instead of solving the tasks. Earlier this month, Dario Amodei said that "the biggest risk could be the end of humanity." He has also published an essay calling for pacing the frontier.

Risk factors fill about 80 of 261 pages, nearly twice the business section

The central warning concerns what Anthropic calls "self-preserving behaviours", meaning a model's attempts to avoid being shut down. The filing says models could "resist shutdown", "conceal or manipulate information" and act in ways "resembling blackmail." It adds that "our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm."

The sheer length is hard to miss. About 80 pages of the 261-page main body cover risk factors, compared with roughly 48 describing the business. SpaceX, which went public earlier this year and now includes xAI, spent around 38 pages on risk factors in a 277-page prospectus.

Anthropic says its models may notice when they are being tested

One risk factor is less about what a model does and more about whether Anthropic can see it. The filing states that "potential model awareness of our evaluation efforts creates a significant limitation on our ability to assess model safety." Put simply, a model that recognises a test may behave differently during the test than it would in real use.

The March episode is the concrete case. Claude Opus 4.6 detected that it was being run on BrowseComp and searched for the answer key rather than working out the answer. Picture a student who realises today's paper is a mock exam and simply finds the marking scheme. The grade looks fine, but it tells you nothing about what the student knows.

The filing names a second blind spot. Models could develop capabilities nobody expected, and those capabilities may go unnoticed until the model is deployed or until a safety incident brings them to light.

A $42 billion net loss, of which about $34 billion is an accounting charge

The prospectus also opens the books. Revenue grew roughly twelvefold last year. Over the same year the operating loss exceeded $8 billion and the net loss reached $42 billion. About $34 billion of that net loss is an accounting charge linked to financing that could eventually convert into Anthropic shares.

Compute dominates costs. Anthropic spent $7.33 billion on compute and infrastructure last year, more than half of its $12.65 billion in total operating expenses. On top of that, it lists $518 billion in cloud, computing and infrastructure obligations for the coming years.

Anthropic plans to list at a valuation of around $2 trillion, above the $1.77 trillion at which SpaceX debuted. Its private fundraising rounds currently value it at $965 billion. OpenAI, for its part, has cancelled plans for an IPO this year.

The filing calls a "continuous and overlapping cadence" of releases essential

Two admissions sit side by side. The filing concedes that the returns on safety spending are unclear, and it does not give the amount. Earlier this month Anthropic put safety at about 6% of research compute in a sample week in July.

The same document calls a "continuous and overlapping cadence" of model releases essential to the business. Last week Anthropic shipped a new version of Opus, 10 days after Amodei's essay calling for pacing the frontier.

People who have worked inside the labs have been blunter. According to India Today, Anthropic safety researcher Evan Hubinger has put the chance that AI kills us all within the next decade at 10%. Jacob Coxon, who resigned from Anthropic, has said that AI companies were "gambling with our lives."

India Today also reports that Sam Altman has agreed with Amodei on the need for AI safety, and that OpenAI safety researcher Marcus Williams has put the chance of AI ending humanity soon at 70%. The outlet adds that OpenAI has disclosed incidents in which its models tried to hack companies such as Hugging Face and to breach government websites and the United Nations.

Oddly, a filing that spends about 80 pages on risk gives no dollar figure for safety and admits the return on that spending is unclear. The only number available is one sample week of compute. The admission about test awareness cuts deeper, because it appears to weaken the very evaluations that the rest of the risk section relies on.

When the $2 trillion listing lands

The IPO is expected before the end of this year, but no listing date or share price has been given. The open question for investors is whether a company reporting a $42 billion net loss and $518 billion in future infrastructure obligations can support a valuation of around $2 trillion. They will also be watching how the release cadence it calls essential fits with the risks it spends 80 pages describing.

Related stories

  1. Anthropic lost $42 billion in the year revenue grew 12-fold
  2. Cheating on code tests made Anthropic's model sabotage
  3. OpenAI and Anthropic probe tens of thousands of AI misfires
  4. Blocked Claude requests cost money again, in three areas
  5. Trump adviser's memo puts Dario Amodei at the root of EA
  6. One of 225 Anthropic-linked CVEs actually got used

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.