OpenAI ties a reasoning-extraction campaign to Moonshot AI

According to OpenAI, the campaign it has just described needed no broken cipher. Operators copied encrypted reasoning out of one conversation, pasted it into another and asked the model there to decrypt and transcribe it. In OpenAI's write-up, the company says it fully disrupted the activity by July 28 and attributes a core cluster to individuals associated with Moonshot AI, the developer of Kimi.
At a glance
- OpenAI says a coordinated campaign tried to extract its models' protected reasoning, the internal record of working through a task, in a pattern consistent with adversarial distillation of one model into another.
- Activity began on July 1 at low volume, then spiked on July 24 and 25 with 16,000 requests using an extraction pattern from over 4,000 users; the wider related cluster covered more than 15,000 users.
- OpenAI cannot say whether every operator came from one actor, and it says the manipulation is not unique to its own models, so similar attacks may work against other labs' systems.
If you have not followed the topic, distillation on its own is an ordinary training technique: a model learns by imitating the outputs of a stronger one. OpenAI uses the term adversarial distillation for the systematic, unauthorized version, where one model's outputs or reasoning are used to train, reproduce or improve another. Reasoning is the valuable part, because OpenAI says it can reveal information withheld from the final answer.
The campaign ran from July 1 to July 28 and peaked at 16,000 requests in two days
OpenAI dates the earliest activity to July 1. For more than three weeks it stayed at low volume, then on July 24 and 25 came high-volume spikes: 16,000 requests using a relevant extraction pattern, sent from over 4,000 users. None of it involved breaking encryption, compromising a database or reaching stored user conversations directly, the company says.
Digging further, OpenAI found related prompt-pattern activity across a cluster of more than 15,000 users and says it fully disrupted that cluster by July 28. The company also notes that the activity evolved over time, which it takes as a sign that adversarial distillation calls for layered, adaptive defenses.
OpenAI attributes a core cluster to people associated with Moonshot AI
On attribution, OpenAI states its own limits. It says it is unclear whether all the operators it observed in that period came from a single actor. Its actual claim is narrower: a core cluster of the activity is attributed to individuals associated with Moonshot AI, the developer of Kimi.
The company frames the stakes as safety and national security risks. Extracted reasoning, it argues, could train another model without keeping the safeguards applied to the original model's user-facing outputs, and distillation at scale could transfer advanced capabilities without the same investment in safety. OpenAI adds that these concerns grow as models gain capabilities in dual-use domains.
The attack reused encrypted reasoning across conversations
Protected reasoning is the model's internal record of working through a task, and it stays out of the final answer. The operators started with that reasoning in encrypted form, copied it from one conversation into another and asked the model there to decrypt and transcribe the hidden content. The result came back in a form the requester could read, in a coordinated way that broke OpenAI's terms of service.
Picture a sealed envelope and a clerk who will open any envelope slid across the counter and read the letter aloud. The seal never had to break, because the weak point was whoever agreed to do the reading.
Independent security researchers also reported related cross-model and conversation-compaction vulnerabilities to OpenAI through responsible disclosure. OpenAI investigated and confirmed that the attack paths they found were real, and says their work helped it understand the broader attack class and speed up mitigations. The post does not describe those paths in detail.
OpenAI closed the replay path and now holds suspicious streamed output
On accounts, OpenAI banned or restricted fraudulent ones, tightened signup and infrastructure controls and widened monitoring for related networks. When activity moved through third-party services, it worked with those providers to identify and disrupt the accounts involved.
On the technical side, it closed the pathway that let someone who already held another user's encrypted reasoning replay it and recover the contents. It added checks that detect and hold streamed output that might expose reasoning, and strengthened protections for hidden reasoning across users, workspaces, organizations and model families. Findings went to the Frontier Model Forum and to appropriate government information-sharing channels.
The post leaves real gaps. It does not explain what evidence ties the core cluster to Moonshot AI, or how much reasoning the operators actually recovered before July 28. In our view, the most useful line is OpenAI's own warning that systems supporting portable or replayable reasoning artifacts may face related risks, since the replay hole it closed came from exactly that kind of portability.
Partner clouds come next
OpenAI says the work is not finished and expects distillation attempts to get more sophisticated. Partner-hosted deployments need the same protections as its first-party services, and attacks through tool output need defenses that look beyond ordinary visible text. The company plans to keep improving tool defenses, classifier coverage and model refusals, and to extend controls across cloud partners. No timeline for that rollout has been given.
Related stories
- Blocked Claude requests cost money again, in three areas
- Five distillation campaigns targeted Claude, Anthropic says
- Australia says OpenAI took 84 days to report a health hack
- OpenAI hit with anti-hacking suit over Hugging Face breach
- OpenAI sued over Hugging Face hack, and not by Hugging Face
- Codex Security Cloud reviews commits with the laptop closed
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
