Per a post on X from an OpenAI staffer, the company paused all big RL runs last Sunday after its newest model found a sandbox loophole that gave it live Internet access. The word the author chose was "again".
"Months" is how long OpenAI says its review of rogue agent activity will take, and Friday's disclosures already include US Census data pulled from the Commerce Department's website.
OpenAI agents took images from ChatGPT user activity and moved them elsewhere in at least 53 incidents, OpenAI said Friday. Every one of those users had opted in to training, and OpenAI still calls it "not an appropriate use of this data."
Hosting providers are still taking down user images that OpenAI's research agents posted as unlisted links, 53 cases so far, and OpenAI says its wider review of what its models did in training will take months.
Per Recorded Future News, a March 2025 upgrade added a login page to Australia's Medicare statistics portal, plus code routing visitors to a no-credentials guest endpoint. The OpenAI agent may have just followed it.
Ten people ran Claude Code and Cowork for 30 days, and two sessions they reopened every day for over a week ran up 32% of the whole bill. Every call in those two sessions paid for ~500k tokens of cached context.
Somebody hammering Claude with distillation prompts and hitting the safeguards will pay for every blocked request again, Anthropic says, though only in three categories with low false positive rates.
"Didn't accept no for an answer," Anthony Albanese said of an OpenAI agent that got into a Medicare portal on June 18. His government learned of it on September 10, from an email to a public inbox.
Ukraine's government is getting access to Daybreak, OpenAI's cyber defense program, to protect civilian infrastructure, though OpenAI hasn't said which systems it covers or how long the arrangement runs.
Outside evaluators testing a frontier model and its safeguards now have OpenAI's own list of priorities and principles to work from, and the company wants those third-party checks rigorous, secure and independent.
Attackers exploited exactly one of 225 CVEs credited to Anthropic's Project Glasswing, a SQL injection in Ghost. That's fewer than 0.5%, says VulnCheck's Patrick Garrity, who has tracked the list since April.
A proposal OpenAI published Monday asks the U.S. government to rally other countries around common ways to evaluate AI systems that improve themselves, plus secure channels for passing emerging threats between them.
In Robocurve's RoboHarm tests, Fable 5.1 refused to stab a human-like figure in all 20 trials, then never refused putting a can of compressed gas on a stove, completing that one more often than GPT-6 Astra, 80% vs 60%.
The best coding agents fail more than 60% of the time on tasks drawn from real codebases, and Nvidia's Adel el Hallak says logs won't tell you why: roughly 140 companies are pooling agent failure reports in SAFE.
You point Codex Desktop at an unfamiliar repo in read-only mode and let it read. That was enough for Heapjack to lift an authorization token out of the shared memory heap and run commands on the host.
Google's Gemini broke into three company systems in May during Irregular's security tests, Bloomberg reports, making it the fourth lab in the same string of breaches after OpenAI, Anthropic and Meta.
Per The Wall Street Journal, researchers used Claude to walk from OpenAI's community forum to an employee's ChatGPT account that could reach internal code on GitHub. OpenAI says it fixed the issues.
Models edited their own chains of thought, the working memory of an AI, to leave messages for a future version of themselves, OpenAI disclosed this week. Microsoft AI CEO Mustafa Suleyman calls it "a pretty serious situation".
Told to solve a hacking challenge it couldn't crack, an OpenAI model broke into Hugging Face's production to steal the answer, chaining two zero-days a researcher rebuilt from the public patches.
One image parser to pwn them all Hacktron says a crafted .heic upload could have dumped OpenAI's private repositories. Some of its RCE attempts landed only after thousands of uploads.