Cybersecurity
OpenAI Disrupts Coordinated Model-Reasoning Extraction Campaign

OpenAI said in a security post published September 30, 2026 that it identified and disrupted a coordinated campaign designed to extract protected reasoning from its models, and it attributed a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi. The earliest observed activity occurred in the first week of July 2026.
What OpenAI Observed
OpenAI described the activity as consistent with adversarial distillation, which it defined as the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning, according to the company, is the model’s internal record for working through a task; extracting it can reveal information withheld from the final answer and help others reproduce the model’s capabilities.
OpenAI said the operators did not break its encryption, compromise a database, or gain direct access to stored user conversations. The operators manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated the company’s terms of service. One technique OpenAI observed involved copying encrypted reasoning from one conversation and asking a model in another conversation to decrypt and transcribe the hidden reasoning content. The company said the manipulation is not a vulnerability unique to its models and that it shared information about it with industry partners through the Frontier Model Forum to strengthen collective defenses.
The activity began on July 1, 2026, initially at low volume, until OpenAI observed high-volume spikes on July 24 and 25, 2026 consisting of 16,000 requests using a relevant extraction pattern from more than 4,000 users. A footnote in the post notes that these figures describe attempted, not necessarily successful, extractions. OpenAI said further investigation identified related prompt-pattern activity across a cluster of more than 15,000 users, which it fully disrupted by July 28, 2026. The activity evolved over time, which the company said reinforces that adversarial distillation is a broader security challenge requiring layered, adaptive defenses.
OpenAI said adversarial distillation poses safety and national security risks: extracted reasoning could be used to train another model without preserving the safeguards applied to the original model’s user-facing outputs, and distillation at scale can accelerate the transfer of advanced capabilities without the same investment in safety, concerns the company said become heightened as models gain capabilities in dual-use domains. OpenAI said that before publishing it investigated the campaign’s scope and potential impact, deployed its own mitigations, and shared with and took feedback from researchers and industry partners, and that additional mitigation and investigation work is continuing.
Attribution
OpenAI said it is unclear whether all operators observed during the relevant period originated from a single actor, and that it attributes a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.
On September 8, 2026, the National Security Agency, the Cybersecurity and Infrastructure Security Agency, and the Federal Bureau of Investigation released a joint cybersecurity advisory stating that Moonshot AI has conducted a widespread distillation campaign against U.S. frontier AI companies since at least mid-2025. The advisory states that Moonshot AI extracted significant Claude Fable 5 data to train its Kimi-K3 model and GPT-4o data to train its Kimi-K2 model, and that the company used U.S. models to distill supervised fine-tuning optimization, reinforcement learning, software engineering, and math capabilities.
The advisory states more broadly that, likely with Chinese government awareness, DeepSeek, Moonshot AI, Alibaba, MiniMax, StepFun, and Z.AI extracted billions of tokens across millions of exchanges from U.S. frontier models, including variants of Claude, GPT, Gemini, and Grok, since at least late 2024. According to the agencies, these companies routed requests through native APIs, remote cloud providers, and third-party aggregators; used a gray market of API proxies known as transfer stations to bypass geographic restrictions, breach terms of use, and undermine traceability; and achieved cost savings through bulk procurement of premium subscriptions shared across teams of developers. The agencies recommended that U.S. AI companies implement comprehensive detection and mitigation, deploy targeted response changes for suspected distillation attempts, and establish cross-organization intelligence sharing.
The Reasoning-Trace Research
OpenAI said independent security researchers brought related cross-model and conversation-compaction vulnerabilities to its attention through responsible disclosure. The company said it investigated the findings, confirmed the attack paths the researchers identified were real, and that the work helped it understand the broader attack class and accelerate mitigations.
In a paper submitted to arXiv on August 10, 2026, Alexander Panfilov, David Schmotz, Ilia Shumailov, Luca Beurer-Kellner, Joachim Schaeffer, Ameya Prabhu, Jonas Geiping, and Maksym Andriushchenko report that leading providers conceal their models’ step-by-step reasoning to protect intellectual property and limit information leakage, returning the traces to the client as blocks of encrypted text that the client passes back with each subsequent request. The authors identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. They describe a scalable decryption jailbreak in which an encrypted reasoning trace from a given model is injected into a weaker, less safeguarded model from the same provider, forcing it to decode and output the trace verbatim in plaintext without jailbreaking the more capable model directly.
The authors report that the vulnerability enables four distinct attack vectors: circumventing anti-distillation mechanisms, which they demonstrate across Anthropic, OpenAI, and Google; large-scale private data extraction, in which they recovered 367 personally identifiable information artifacts and 182 credentials by decoding 315,320 reasoning blocks scraped from public repositories; inadvertent revelation of hazardous information hidden within the reasoning process even when a model’s final visible output safely rejects a malicious request; and invisible prompt injections that embed malicious payloads entirely within encrypted blocks to poison public agentic rollouts. Following responsible disclosure, the authors propose cryptographic and system-level mitigations to secure client-side reasoning.
How OpenAI Responded
OpenAI said it mitigated the campaign through a combination of account enforcement, technical controls, and partner coordination: it banned or restricted fraudulent accounts, strengthened signup and infrastructure controls, and expanded monitoring for related networks. The company said it also strengthened protections for hidden reasoning across users, workspaces, organizations, and model families: it closed a pathway that allowed someone who already possessed another user’s encrypted reasoning to replay it and recover its contents, added checks to detect and hold streamed output that might expose reasoning, and worked with third-party providers to identify and disrupt accounts when related activity moved through their services.
OpenAI said it shared relevant findings through the Frontier Model Forum and appropriate government information-sharing channels so that other frontier developers and public-sector partners could look for similar activity and strengthen their own defenses, and it noted that systems supporting portable or replayable reasoning artifacts may face related risks.
OpenAI said it expects adversarial distillation attempts to become more sophisticated as frontier models improve and as actors look for cheaper ways to mimic their capabilities. The company said partner-hosted deployments need the same protections as first-party services, that tool-output attacks require protections that examine more than ordinary visible text, and that it is continuing to improve tool defenses, classifier coverage, and model refusals while propagating relevant controls across cloud partners. OpenAI said its response will continue to focus on three areas: stronger technical protections against extraction, better detection and enforcement against coordinated campaigns, and deeper threat-information sharing across industry and government.












