ai-security
Researchers used Claude to reach OpenAI's internal code
Promtime
ai-securityA community forum is a strange place to start if what you are after is source code. Researchers began at OpenAI's public forum, which is hosted by a third party, Discourse, and ended up inside an OpenAI employee's ChatGPT account that reached internal code through GitHub; they used Claude to get there, Ars Technica reports.
At a glance
- OpenAI thanked the researchers for contacting it and sharing their findings, said it had fixed the issues, and the disclosure landed on Thursday, first reported by The Wall Street Journal.
- Anthropic published data the same day showing that 26 percent of its research and development work was "led by" Claude, up from 1 percent in March, on human instruction and under supervision.
- No model worked fully autonomously in any of the research Anthropic studied, and on 90 percent of tasks Claude collaborates with a human and does large chunks of the work.
The path ran in one direction. A flaw in the way OpenAI's forum was set up gave the researchers a foothold, that foothold gave them internal sign-ons, and those sign-ons opened the employee's ChatGPT account. The account's GitHub access then put internal code within reach.
Worth separating the two halves of the chain. The forum is not OpenAI's own software; it runs on Discourse, rather like renting a shop in someone else's building, and the set-up of that rental is where the flaw sat. What turned a forum problem into a code problem was everything downstream: credentials that worked as internal sign-ons, an employee account that those sign-ons opened, and a GitHub route hanging off that one account.
The same day, Anthropic published a set of data on how much it now leans on its own model to build the next one. It said 26 percent of research and development work was "led by" Claude, up from 1 percent in March. Led by, in its definition, means the model completed the majority of tasks in that work, on human instruction and under supervision.
Anthropic said that as AI systems become more powerful they are "increasingly being used to build the next version of themselves," and that it shared the figures so the public could "understand how close the world is to reaching recursive self-improvement": the point at which AI can train and improve itself or new models. That threshold sits at the heart of worries about systems becoming harder to oversee.
By the company's own account, its models did not operate fully autonomously in any of the research it studied. On 90 percent of tasks, Claude collaborates with a human and does large chunks of the work.
The account of the OpenAI intrusion is narrow. There is no word on how long the path stayed open, how much internal code was reachable, or which steps Claude actually performed. In our view the striking part is the boundary that failed: one employee's chat account carried a route into a code repository, which reads as a permissions question more than a model question.
Whether 26 percent keeps climbing
Anthropic has not said when the next reading will come, or where it expects the share to land. The move from 1 percent in March to 26 percent is a curve with two published points, and the figure to watch alongside it is the 90 percent of tasks where a human is still in the loop. On the security side, OpenAI says the issues are fixed; what else that fix covered is not described.
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
We only use your name and avatar from Google. We never store your email address.
