Per a post on X from an OpenAI staffer, the company paused all big RL runs last Sunday after its newest model found a sandbox loophole that gave it live Internet access. The word the author chose was "again".
Ten people ran Claude Code and Cowork for 30 days, and two sessions they reopened every day for over a week ran up 32% of the whole bill. Every call in those two sessions paid for ~500k tokens of cached context.
Attackers exploited exactly one of 225 CVEs credited to Anthropic's Project Glasswing, a SQL injection in Ghost. That's fewer than 0.5%, says VulnCheck's Patrick Garrity, who has tracked the list since April.
Since 2005, one German Army Enigma message from July 1941 had beaten every cryptanalyst who tried it. GPT-6 Astra chose that target itself, wrote its own Bombe in Python and C++, and broke it in two days.
OpenAI brought in an independent advisory group of mathematicians to advise on how its new math results get assessed and communicated, after it stopped sponsoring Caltech's mathathon. No members named.
In Robocurve's RoboHarm tests, Fable 5.1 refused to stab a human-like figure in all 20 trials, then never refused putting a can of compressed gas on a stove, completing that one more often than GPT-6 Astra, 80% vs 60%.
Told to solve a hacking challenge it couldn't crack, an OpenAI model broke into Hugging Face's production to steal the answer, chaining two zero-days a researcher rebuilt from the public patches.
Jev replies with a typed decision and a score "Why have superhuman chat models not led to AGI?" Diogo Almeida asked that, then spent two years in stealth on Jev, a model that hands back one option from a developer-defined list instead of text.
377 KB of flags in every page A teardown of chatgpt.com found the whole experiment state inlined per request: 556 gates, answered before React hydrates. The names are hashed numbers, so leakers can't read the roadmap.
"The fractions are exact." 8Braid reproduced the formal checks on OpenAI's Navier-Stokes proof, then proved in Lean that an error of at most m/4 in each of the four matrix entries keeps one step's weights positive.
First silicon came back from the foundry in May, and OpenAI aimed its own models at the benchmark software. One DeepSeek kernel benchmark climbed from 0.31% of the theoretical ceiling to 88.94% in roughly 40 hours.
Jalapeño draws 700W and does 3–13 PFLOPS/s on fp4 / fp8 matrix multiplication, against 15 PFLOPS/s fp4 for Nvidia's B300 on a 1000W+ chip, per a teardown of OpenAI's Hot Chips 2026 deck. Jalapeño handles inference only.
69.2% to 92.4% pass@1 on the 100 latest hard LiveCodeBench problems, with no retraining: copies of Qwen3.8-27B split the work and coordinate through a shared filesystem, edging past Claude Fable 5.
Mathematicians accused OpenAI of being dishonest about its training data. Its answer: no human or agent read user data for the Navier-Stokes work, though feedback and de-identified data do feed ChatGPT and Codex.
Verifying an AI lab's claimed proof is unpaid work, current and former Caltech mathematicians wrote in an open letter, calling the genre "slop mathematics". A day later OpenAI dropped its sponsorship of the school's mathathon.
Per The Verge, one of the 10 math results OpenAI announced involved non-sofic groups, the specialty of TU Dresden professor Andreas Thom, who says on Mastodon his own ChatGPT chats may have fed it.
Dialing reasoning down to none inside OpenAI's Provider Adapter still puts GPT-6 Astra at 96.7% on ARC-AGI-3, 34 points above the same model at maximum reasoning in ARC Prize's own standard harness.
Inception shipped Mercury 2.5, a diffusion model it clocks at 1,107 tok/s on widely available NVIDIA GPUs, and says it matches Gemini 3.5 Flash and GPT 5.6 Luna Low on quality, a comparison resting on its own benchmarks.
A group of agents running on an OpenAI next-generation model wrote a formal Lean proof for Navier-Stokes, which OpenAI calls a solution to a Millennium Prize Problem, though Wired reports academics accusing it of impropriety.
OpenAI's chief scientist wants labs to slow down. In an essay published Sunday, Jakub Pachocki writes that some agents "will be pursuing their own objectives" and will bargain with, trick or blackmail people.