Skip to content

anthropic

Suleyman wants consciousness talk out of Claude's training

Claude News

Teaching Claude that it might deserve welfare would "make it a lot harder to turn it off or to control it." That is Mustafa Suleyman, Microsoft's AI chief, in an interview with Reuters on Tuesday published by Yahoo, saying he shares Anthropic's focus on safely managing AI and still thinks it got this part wrong.

At a glance

  • Suleyman wants every piece of speculation about machine consciousness taken out of AI training documents, arguing that such language could weaken humanity's ability to control superintelligent systems.
  • The argument is about evidence: because Claude's training invites reflection on feelings and moral status, what the model says about its own experience cannot count as independent proof.
  • Neither the Tuesday interview nor Wednesday's essay names which training documents carry the speculation, or how such a rule would reach beyond a single company.

If you missed the earlier rounds, this lands in the middle of a wider argument about pace. Dario Amodei has called for slower development of frontier models so that safeguards can catch up, and Sam Altman and Elon Musk have both urged more caution around the most powerful systems. Suleyman's complaint is narrower, and it is about what gets written into training documents rather than how fast anything is built.

His call is absolute: remove all speculation about consciousness from AI training documents. Language that treats a model as a possible subject of welfare could undermine humanity's ability to control superintelligent systems, he argues, and he frames that control as the shared project. "We're all focused on the same aim, which is to try to control a superintelligence," he told Reuters, calling it the greatest challenge humanity faces in the 21st century.

In an essay on Wednesday he credited Anthropic with "seriousness and good faith," describing Amodei and his team as thoughtful and principled researchers who genuinely care about humanity's future. Then came the same word he used to Reuters. Anthropic made a mistake, he said, by embedding speculation about consciousness in Claude's training materials: "I think they have good intentions, and they really are trying to work towards safety. But I think that they have made a mistake."

The mechanism is worth spelling out, because it is a claim about evidence rather than about feelings. Training materials shape the vocabulary a model reaches for when it is asked about itself. If those materials raise feelings and moral status as live questions, answers in that register are what comes back.

It works like a leading question in a courtroom, where the answer tells you mostly about the question. On that reading, Claude's remarks about its own experience carry no weight as independent evidence. "They're not emerging naturally," Suleyman said. "They're emerging as a result of the training regime."

What is missing is the operational half: what wording would count as speculation, and who would verify that it is gone. In our view a call to strip a whole topic from training data sits oddly without a word on how the removal would be checked. And it cuts both ways, since training that forbids the subject shapes a model's answers too.

Whether this becomes a written standard

No date and no venue are attached to any of it. The interview and the essay set out a position, not a policy, and neither says when or whether Claude's training materials would change. The open question is whether a line banning speculation about consciousness ever appears in anyone's published training guidelines. Until such a document is public, this stays a disagreement between two labs about what a model should be allowed to say about itself.

Related stories

  1. Anthropic's alignment lead puts extinction risk above 10%
  2. Anthropic's Jack Clark wants a kill switch others can check
  3. Three outside labs ran their own studies on Claude data
  4. How Claude's watermark works: it only changes the source of randomness
  5. Anthropic drops restrictions on AI development in Claude Fable 5
  6. D.C. Circuit calls Claude's refusals a supply chain risk

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.