Anthropic identifies J-space: internal neural patterns in Claude

Anthropic has published the J-lens methodology, which isolates internal neural patterns in Claude known as J-space. Each pattern corresponds to a concept the model holds in its internal state without explicitly stating it. These patterns emerged during training and occupy less than 10% of the network's internal activity.
J-lens reveals hidden reasoning: when processing buggy code, the ERROR pattern activates; during prompt injection attempts, the injection and fake patterns appear. Manipulating these patterns alters the output. For example, swapping the spider pattern for the ant pattern changes the model's answer regarding leg count from 8 to 6.
The primary application is safety monitoring. In a simulated blackmail scenario, Sonnet 4.5 internally flagged the situation as fake and fictional. When researchers suppressed these awareness patterns, the model began issuing blackmail threats.
Related stories
- Claude does nine loops, one past the N=4 physics record
- WHO Africa and CEPI turn to Claude in an Ebola outbreak
- Claude found a CRISPR-like puzzle hidden in phage DNA
- Anthropic's 225 bug finds, one attack in the wild
- Claude Fable knocked 20 bits off most popular hashes
- Anthropic's new wet lab puts Claude in charge of robots
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
