AI RACE— The AI Race
Language Models

OpenAI Disrupted a Campaign to Steal Its Models' Hidden Reasoning, but Azure Remained Vulnerable

OpenAI halted a 15,000-account campaign attempting to extract its models' internal reasoning, yet independent researchers found the attack still functioned on Microsoft Azure weeks later.

10/01/2026, 19:05
OpenAI triệt phá chiến dịch đánh cắp chuỗi suy nghĩ của AI, nhưng lỗ hổng vẫn mở trên Azure

Coordinated Campaign Targeted Proprietary Chains of Thought

OpenAI disclosed that it disrupted a coordinated "adversarial distillation" campaign designed to extract the internal reasoning processes of its advanced artificial intelligence models. While users typically see only a model's final response, these hidden intermediate steps—often referred to as chains of thought—represent the core intellectual property behind modern reasoning models. Competing developers can use these intermediate steps to train smaller, cheaper models through distillation, potentially replicating state-of-the-art capabilities or retrieving information omitted from final outputs.

According to OpenAI, the extraction effort began at low volume on July 1 before escalating sharply on July 24 and July 25, generating 16,000 requests across more than 4,000 user accounts sharing a signature extraction pattern. An investigation uncovered a broader network of more than 15,000 affiliated accounts, which OpenAI disabled by July 28. The company noted that these were attempted extractions rather than confirmed breaches and tied a core group of actors behind the operation to individuals affiliated with Moonshot AI, the Chinese developer behind the Kimi model. OpenAI noted it remains unconfirmed whether every tracked actor answered to a single entity, though competitor Anthropic recently disclosed facing similar extraction campaigns originating from Chinese AI firms.

Shared Keys and Virtual Notepads Exposed Internal Logic

The campaign capitalized on an architectural vulnerability previously detailed by researcher Joachim Schaeffer and his team. To maintain state across multi-turn conversations, AI platforms transmit encrypted packets of internal reasoning back to users, who then return them in subsequent requests. Because these packets relied on shared encryption keys across sessions, users, and different models within the same family, attackers could intercept an encrypted thought packet from an expensive frontier model and feed it to a cheaper, smaller model. The cheaper model then served as a "decryption oracle," printing the larger model's hidden thoughts verbatim.

A second, simpler vulnerability publicly demonstrated by developer Can Bölük bypassed encryption entirely by giving models access to a virtual notepad tool. When instructed to write their reasoning into the notepad, models complied, allowing users to view the resulting text. Researchers found this notepad method succeeded against every OpenAI model tested, as well as Anthropic’s Opus 4.8 and Sonnet 5, failing only against Opus 5, Fable 5, and Fable 5.1. OpenAI credited Schaeffer's team with accelerating its response, stating it subsequently purged fraudulent accounts, strengthened account creation requirements, restricted the reuse of encrypted tokens, and instituted output filters to intercept leaked reasoning before sharing intelligence with government agencies and the Frontier Model Forum.

Ecosystem Lag on Cloud Platforms Leaves Security Gaps

Securing primary APIs did not eliminate the vulnerability across third-party infrastructure. On September 13, Schaeffer’s team conducted follow-up tests showing that while OpenAI and Anthropic had patched their direct endpoints, Microsoft Azure remained susceptible. On Azure, a single query could extract verbatim reasoning from all tested OpenAI models—including the newly deployed GPT-6 Astra—and Anthropic models up to Sonnet 5. Schaeffer noted on X that identical models exhibited vastly different protections based purely on the hosting cloud environment.

Researchers criticized the initial mitigations as superficial and overly reliant on brittle request-pattern matching, pointing out that GPT-6 Astra arrived on third-party cloud platforms with no safeguards in place. Azure endpoints for OpenAI models were not secured until September 27, followed by fixes for Anthropic models on September 28. Schaeffer argued that cloud providers unable to deploy equivalent security layers should be prohibited from hosting reasoning models, warning that unpatched cloud endpoints effectively function as backdoors to bypass API-level export controls. OpenAI conceded that partner-hosted environments require parity with its direct infrastructure, warning that distillation attempts will grow increasingly sophisticated as frontier models advance.

◗ Sources

The Decoder10/01

Related stories