Disrupting a Coordinated Model-Distillation Campaign
We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld from the final answer and help others reproduce the model’s capabilities.

TL;DR
- OpenAI identified and disrupted a coordinated campaign to extract protected reasoning from its models, occurring since early July.
- The activity, termed adversarial distillation, involves unauthorized use of model outputs or reasoning to train other models.
- Operators manipulated model interactions to reproduce protected reasoning, not by breaking encryption or databases.
- The campaign involved significant activity on July 24-25, with over 16,000 requests from 4,000+ users, and was fully disrupted by July 28.
- A core cluster of activity is attributed to individuals associated with Moonshot AI, the developer of Kimi.
- Adversarial distillation poses safety and national security risks by enabling the transfer of advanced capabilities without safeguards.
- OpenAI responded with account enforcement, technical controls, partner coordination, and strengthened reasoning protections.
- The company is continuing to improve defenses, detection, and information sharing to counter increasingly sophisticated attempts at adversarial distillation.