Story
October 1, 2026
OpenAI Says Hidden Reasoning Became the New Target in AI’s Security Race
OpenAI says it halted a July campaign that tried to extract hidden model reasoning at scale, tracing a core cluster to people associated with Moonshot AI. The company argues the episode exposes an industry-wide race to stop rivals from cheaply reproducing frontier capabilities.
OpenAI says the campaign began quietly on July 1, when operators started probing ways to pull “protected reasoning” — the model’s internal working record — into visible responses. The company describes the practice as adversarial distillation: using a model’s outputs or reasoning to train, reproduce or improve another system without authorization.1
The operation then accelerated. OpenAI says it saw high-volume spikes on July 24 and 25, with 16,000 extraction-pattern requests from more than 4,000 users. Its investigation ultimately identified related activity across a cluster of more than 15,000 users, which it says it fully disrupted by July 28.1
According to OpenAI, the operators did not crack encryption, breach a database or obtain stored conversations directly. Instead, they allegedly manipulated interactions — in one method, copying encrypted reasoning from one chat and asking a model in another to decrypt and transcribe it. Independent researchers had separately flagged related cross-model and conversation-compaction weaknesses, OpenAI said, helping it map the broader attack class.1
The company says it closed the replay pathway, tightened output checks, banned or restricted fraudulent accounts and coordinated with third-party providers. It also shared findings with the Frontier Model Forum and government channels, arguing that portable or replayable reasoning artifacts could create similar risks across advanced AI systems.1
Attribution is the most politically charged part of the disclosure. OpenAI says it cannot determine whether every operator came from one actor, but attributes a “core cluster” to individuals associated with Moonshot AI, the Chinese developer behind Kimi. A report on the announcement notes that Anthropic has made similar allegations involving Moonshot.2
OpenAI’s broader warning is that stolen reasoning could transfer advanced capabilities without carrying over the safety restrictions imposed on public-facing answers. As models become more capable, it says, attempts to turn hidden reasoning into a shortcut will only grow more sophisticated.1