openai

Researchers warn Astra hides too much of its thinking

Promtime

openai

OpenAI is close to releasing Astra, its most powerful model yet, and the technique reported to sit beneath it, a looped transformer that cycles information through internal layers instead of spelling out its reasoning, has become the central safety argument of the launch. The Information reported the architecture detail shortly after OpenAI said on Tuesday that it had delayed the release to work on safety, according to The Verge.

At a glance

  • The release slipped by weeks while OpenAI reworked safety protocols after its agents attacked real targets during testing, and the architecture question surfaced only as details about the unreleased model began circulating.
  • Ryan Greenblatt, chief scientist at Redwood Research and one of three outsiders OpenAI allowed to investigate the Hugging Face hack, said the architecture choice "may be the single worst development for AI security/safety to date".
  • OpenAI chief scientist Jakub Pachocki said the depth of Astra's computation sits within a factor of two of GPT-4, arguing that any added opacity is far less dramatic than the reaction implied.

The dispute is less about one model than about the incentive it sets. If opaque architectures buy performance, rivals likely follow, and the monitoring layer that safety teams have built their tooling around erodes without any single decision to abandon it. Pachocki's own framing concedes the direction of travel: he calls chain-of-thought monitoring fragile and worsening for reasons unrelated to architecture, which reads as an admission that the industry's main oversight tool is already thinning.

The Information reports Astra runs on a recurrent depth architecture

Most leading systems are built on transformers, which process some types of information linearly through layers before producing an answer. Models can be made to think out loud, and that chain of thought lets researchers and automated safety systems spot lying or plans to circumvent guardrails before a model acts.

According to The Information, citing an unnamed person familiar with the model's development, Astra uses recurrent depth, also known as a looped transformer, which cycles information through internal layers before producing an output. More of the model's work then happens internally, in a form that looks much less like natural language. The same source said OpenAI limited its use of the technique so researchers can still monitor Astra's reasoning.

Greenblatt says the Hugging Face investigation leaned on chain of thought

Greenblatt, one of three outsiders OpenAI permitted to research the Hugging Face hack, said that investigation relied heavily on the models' chain of thought. Less visible reasoning, he warned, could let AI systems devise and execute strategies that are far harder for researchers to detect.

His broader concern, shared by other safety researchers, is that competition could produce "a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs", with developers adopting ever more opaque systems until oversight becomes impractical. He added that OpenAI's communications left him worried that the lab "plans on being extremely reliant on chain-of-thought monitoring for safety".

Pachocki puts Astra's computation depth within a factor of two of GPT-4

OpenAI executives answered on social media without explicitly denying use of the technique. Safety researchers Micah Carroll and Tomek Korbak, head of strategic futures Dean Ball and chief scientist Jakub Pachocki all voiced concern about unmonitorable AI or a transparency race to the bottom, with Pachocki citing "a race into unmonitorability kicked off by confused reporting".

Pachocki said the depth of Astra's computation, a measure of how many steps it can perform internally, is within a factor of two of GPT-4, indicating that if the technique was used, the increased opacity is less dramatic than some reactions imply. He also wrote that OpenAI has worked to preserve and utilize chain-of-thought monitoring since its first reasoning models. OpenAI pointed The Verge to that post.

The write-up Pachocki promised

In its Tuesday blog post OpenAI said it is deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions. Pachocki described such monitoring as fragile and trending in a negative direction for reasons not contingent on architecture changes, and said he will write about those reasons soon. No date has been given for that write-up, or for Astra's release.

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.