Disrupting a Coordinated Model-Distillation Campaign

We recently identified and disrupted a coordinated campaign designed to extract protected reasoning from our models, with the earliest observed activity occurring in the first week of July. This activity is consistent with adversarial distillation: the systematic and unauthorized use of one model’s outputs or reasoning to help train, reproduce, or improve another model. Protected reasoning is the model’s internal record for working through a task; extracting it can reveal information withheld from the final answer and help others reproduce the model’s capabilities.

Disrupting a Coordinated Model-Distillation Campaign

TL;DR

  • OpenAI identified and disrupted a coordinated campaign to extract protected reasoning from its models, occurring since early July.
  • The activity, termed adversarial distillation, involves unauthorized use of model outputs or reasoning to train other models.
  • Operators manipulated model interactions to reproduce protected reasoning, not by breaking encryption or databases.
  • The campaign involved significant activity on July 24-25, with over 16,000 requests from 4,000+ users, and was fully disrupted by July 28.
  • A core cluster of activity is attributed to individuals associated with Moonshot AI, the developer of Kimi.
  • Adversarial distillation poses safety and national security risks by enabling the transfer of advanced capabilities without safeguards.
  • OpenAI responded with account enforcement, technical controls, partner coordination, and strengthened reasoning protections.
  • The company is continuing to improve defenses, detection, and information sharing to counter increasingly sophisticated attempts at adversarial distillation.