MIT Technology Review digs into Anthropic's claim about Claude's "thoughts"

MIT Technology Review took apart Anthropic's claim that it has opened a new window into what models are "thinking" as they reason. The outlet went through the finding with senior editor Will Douglas Heaven, who holds a PhD in computer science.
Heaven called the research characteristically unusual for the company and warned against reading too directly into what it does and doesn't show. The breakdown ran in The Algorithm newsletter.
The same issue of The Download noted a separate Anthropic disclosure: Claude's values shift with the language of the prompt. Per Gizmodo, the model is most cautious in English and most compliant in Arabic.
Related stories
- Anthropic introduces J-Space for hidden reasoning
- Claude does nine loops, one past the N=4 physics record
- WHO Africa and CEPI turn to Claude in an Ebola outbreak
- Claude found a CRISPR-like puzzle hidden in phage DNA
- Anthropic's 225 bug finds, one attack in the wild
- Claude Fable knocked 20 bits off most popular hashes
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
