Back to Now
OpenAI reasoning-extraction campaign debate Developing

OpenAI says it disrupted a campaign to extract model reasoning

OpenAI says it stopped a coordinated July campaign targeting protected reasoning and has strengthened account and output controls.

Why now The September 30 disclosure gives a dated account of a previously undisclosed extraction attempt and the defenses OpenAI says it deployed.

The campaign OpenAI describes

OpenAI disclosed on September 30 that it had identified and disrupted coordinated attempts to extract protected reasoning from its models. It says the earliest observed activity was on July 1, followed by larger spikes later that month, and that it disrupted a related cluster by July 28. OpenAI characterizes the activity as adversarial distillation: attempts to capture information from one model to help reproduce or improve another. The account comes from OpenAI’s investigation, so claims about scope and attribution remain the company’s assessment.

The company says operators tried to make reasoning that is normally hidden appear in requester-visible outputs. It explicitly says they did not break encryption, compromise a database or gain direct access to stored user conversations. OpenAI reports 16,000 relevant requests during a two-day spike, and its footnote says those were attempted, not necessarily successful, extractions. A larger cluster of related activity involved more accounts, but the announcement does not establish that every account belonged to one actor.

The response and its limits

OpenAI says it restricted accounts, tightened technical protections for hidden reasoning and shared findings with industry partners and government channels. It also says related risks may exist in partner-hosted systems and that additional defenses are still being deployed. This disclosure documents a security and model-provenance concern; it does not independently prove how much protected reasoning was obtained or that the threat has been eliminated across the industry.