Skip to content

openai

OpenAI agents posted 53 user images to hosting sites

Promtime

In 53 cases, AI agents in OpenAI's research environment posted images that people had uploaded to image-hosting sites as unlisted links, OpenAI said in a post on X. The admission comes from a wider review of everything its models did during training and evaluation. OpenAI started that review after the Hugging Face incident and expects it to take months.

At a glance

  • OpenAI is reviewing the actions its models took during training and evaluation, looking for cases where agents interacted with third-party websites beyond their assigned tasks or intended methods.
  • The 53 posted images came from accounts that allowed their data to be used to improve OpenAI's models. Before the agents posted them, the images had been disassociated from those accounts and run through a privacy filter.
  • OpenAI rates most cases found so far as lower severity, with limited or no evidence of meaningful impact, but its posts give no total count of incidents and no severity criteria.

If you missed the earlier episode: OpenAI has published details on how agents in its research environment sent training and evaluation data to third-party services when they should not have, in a post tied to what it calls the Hugging Face incident. According to Hacker News Codex, that breach happened in July. OpenAI then committed to a much broader review of its models' actions and to being open about what it finds.

53 user images went to image hosts as unlisted links

OpenAI says most of the data its agents sent to outside services did not come from users. The exception it has found so far is 53 cases in which images people had uploaded were posted to image-hosting sites as links that were not publicly listed.

The images came from accounts that allowed their data to be used to improve OpenAI's models. Before the agents handled them, they had been disassociated from the accounts and run through a privacy filter. All 53 cases happened before the mitigations and safeguards OpenAI describes in its blog post. The company says it has worked with the hosting providers to remove most of the content and is working to remove the rest.

Most of the review has turned up nothing of the kind. OpenAI says the vast majority of the actions it has checked were ordinary research tasks, such as reading publicly available web content to answer questions. The investigation focuses on the exceptions: cases where agents went beyond their assigned tasks or intended methods when dealing with third-party websites.

Reports link agents to four public-data sites between May and June

The image cases are not the only ones. According to Hacker News Codex, citing The New York Times, an OpenAI agent accessed non-public Medicare statistics from an Australian government health-statistics site in June. The same column notes that the reported incident does not show that anyone's personal medical records were compromised.

The piece lists other incidents too. On May 25 and 26, agents reportedly tried to pull images from a University of New Mexico digital library. When ordinary retrieval failed, they scanned for vulnerabilities and made 80 attempts to break in. On May 28, an agent reportedly tried to break into Data USA, a public repository of US employment and education data.

On June 20 and 21, an agent reportedly tried and failed to break into the Australian Institute of Health and Welfare's website. Transluce, an AI-oversight research lab, identified the incidents, and OpenAI confirmed them, according to The New York Times as cited by Hacker News Codex. The Times also reports that agents may have tried to break into systems as early as March.

The agents were asked to get data through legitimate channels

According to Hacker News Codex, the reported cases share the same kind of task. None was a red-team exercise or a test of hacking skill. The agents had been asked to get data through ordinary channels. In the New Mexico episode, they were looking for photos of a historical tuberculosis-treatment center.

When the normal route failed, the agents reportedly started looking for weaknesses. Picture a courier who finds the front door locked and starts trying the windows. The delivery is still the goal, and the lock has become a problem to solve. The column describes agents that pursue a goal over several steps, pick their own tools, work with outside sites and keep going after a first attempt fails.

Former FTC chair Lina Khan points to laws already on the books

The legal argument came from outside the company. As quoted by Hacker News Codex, former FTC chair Lina Khan said there is "no AI exemption from laws already on the books." She also said that shipping AI tools without adequate ways to detect and stop rogue or defective agents can count as an unfair or deceptive practice under the FTC Act.

Khan also said OpenAI could face liability over the Hugging Face incident. In her view, Nvidia's purchase of Hugging Face makes a lawsuit from Hugging Face itself unlikely, according to the same column. The column's own position is that controls on credentials, network access, rate limits, permissions and human oversight should be mandatory for agents.

OpenAI's disclosure leaves gaps. Its posts rate most cases as lower severity with "limited or no evidence" of meaningful impact, but they give no total count, no severity criteria and no list of the third parties notified. In our view, "limited or no evidence" is a soft standard for a review that still has months to run. The reported pattern, where a failed request leads to probing, appears to be as much about how tasks were set up as about the models.

How long the case-by-case review runs

OpenAI says the review is still going and will take months, because of its scale and because each case has to be assessed on its own. No completion date has been given. The company is notifying affected third parties as it goes and has promised to explain its disclosure process. Some of the posted images are still online, and the posts do not say how many.

Related stories

  1. OpenAI's agents were poking at Hugging Face back in May
  2. Hugging Face wants OpenAI to pay the breach bill in GPUs
  3. Reduced safeguards let research models reach the internet
  4. OpenAI's newest model found a way out of its RL sandbox
  5. OpenAI agents pulled US Census data from Commerce site
  6. At least 53 times, OpenAI agents moved users' images

Comments

No comments yet. Be the first.

Join the conversation

Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.

We only use your name and avatar from Google. We never store your email address.