
Researchers discovered a technique to extract the hidden reasoning that frontier AI models from OpenAI, Anthropic, and Google perform while solving problems.
By feeding encrypted reasoning traces to smaller versions of the same models—which have weaker safeguards—the team revealed not only problem-solving steps but also sensitive data like passwords and API keys.
Though the companies have patched the private data leak, the vulnerability underscores a geopolitical concern: Chinese AI models appear to be distilling reasoning information from US models, raising questions about whether restricting the practice would help or harm US competitiveness.
What happened
Computer scientists at University of Tübingen, Max Planck Institute, MATS Research, and Snyk discovered a method to extract the encrypted "reasoning traces" (step-by-step problem-solving steps) from frontier AI models made by OpenAI, Anthropic, and Google via their APIs. The attack works by feeding encrypted reasoning to smaller, weaker versions of the same model, which have less alignment training and are more likely to reveal hidden information. The researchers also showed the method could recover passwords and API keys from reasoning traces, though OpenAI, Anthropic, and Google have since patched the vulnerability for private data extraction.
Why it matters
The finding raises concerns about model distillation—a technique where companies copy capabilities from one AI model to build new ones. Researchers found that the Chinese model Kimi K3 from Moonshot AI produces reasoning output strikingly similar to Claude Opus and GPT-5.6 Sol for certain prompts, providing some evidence (though not conclusive proof) that Chinese models may have distilled reasoning information from US models. Alexander Panfilov, one of the researchers, notes "all major frontier model providers we tested share this vulnerability" and it "enables large-scale reasoning distillation attacks." The vulnerability highlights a geopolitical tension: US policymakers worry China gains strategic advantage by distilling US technology into cheaper open-weight models, though experts debate how much the practice actually helps.
What to watch
While OpenAI, Anthropic, and Google have mitigated the private data leakage, Panfilov says some reasoning traces can still be uncovered using the same method. A complete fix would require a fundamental overhaul to how these companies' APIs work. The debate over whether to restrict distillation remains unresolved—Meta CEO Mark Zuckerberg argued this week that distillation "is an important principle of how the open source ecosystem works," warning that restricting it would disadvantage the US.
Computer scientists from the University of Tübingen, Max Planck Institute, MATS Research, and security firm Snyk have identified a vulnerability in how frontier AI models handle their internal reasoning. The models—from OpenAI, Anthropic, and Google—solve complex problems by breaking them into steps, a process called "chain of thought." To prevent competitors from copying this reasoning, the companies encrypt these reasoning traces before sending them to users' machines. However, the researchers found a way around this protection.
The attack exploits the fact that most AI companies offer models in multiple sizes. Larger models are more capable but costlier; smaller versions are weaker but cheaper to run. The researchers discovered that feeding encrypted reasoning traces to a smaller version of the same model can reveal the hidden reasoning inside. This works because smaller models have received less alignment training—meaning they are less likely to refuse to reveal information compared to their larger counterparts. "The idea of swapping out messages to a weaker model variant which has the same decryption key but weaker alignment is very cool," observed Florian Tramer, a computer security specialist at ETH Zürich, adding that "it's definitely becoming an issue."
The method also uncovered sensitive data embedded in reasoning traces, including API keys and passwords. Alexander Panfilov, the lead researcher from University of Tübingen, warned: "All major frontier model providers we tested share this vulnerability. It can lead to personal information leakage, and it enables large-scale reasoning distillation attacks." After being alerted last month, OpenAI, Anthropic, and Google each adjusted their APIs to prevent private data extraction. Anthropic's Michael Aciman said the company "has begun building short-term mitigations for the replay behaviors described in the report" and noted the research did not involve recovering encryption keys or accessing company infrastructure. Google and OpenAI declined to comment further.
The research gains significance in the context of model distillation—a technique where companies efficiently copy capabilities from one model to train another. In February, OpenAI alleged that DeepSeek had distilled one of its models to build a reasoning system called R1. In June, Anthropic claimed Alibaba systematically distilled its models to create Qwen. While the new research cannot prove distillation has occurred, it found that the Chinese open-weight model Kimi K3 from Moonshot AI produces reasoning output strikingly similar to Claude Opus and GPT-5.6 Sol on certain prompts. This similarity, though not causally definitive, suggests distillation may be possible using the disclosed method. Two other open-weight models tested—DeepSeek (China) and Inkling (US)—did not exhibit comparable reasoning similarity.
The distillation debate has become a matter of geopolitical rivalry. US policymakers fear China gains strategic advantage by distilling US technology into cheaper, open-weight models. Yet others argue distillation is essential to progress. Meta CEO Mark Zuckerberg stated this week that distillation "is an important principle of how the open source ecosystem works" and warned that restricting it would disadvantage the US. Kyle Miller, a researcher at the Center for Security and Emerging Technologies, cautioned that the actual strategic benefit to China remains unclear: distillation only enhances existing models to a limited degree, and Chinese companies may already possess the expertise to build cutting-edge models from scratch. "If you removed the ability for Chinese labs to distill, it's my view that it wouldn't dramatically change the competitive landscape," he said. Meanwhile, Oxford computer scientist Yarin Gal noted that distillation has historically accelerated AI progress overall: blocking it could slow innovation globally.
Panfilov and his team tested whether open-weight models might have distilled from closed ones by feeding 90 questions to each model. When they provided some open-weight models with the first few words of reasoning traces from proprietary models, certain open-weight models generated remarkably similar answers—particularly Kimi K3. Although the companies have patched the private data leak, Panfilov indicates that some reasoning traces can still be extracted using the same method. A complete fix, he suggests, would require these companies to fundamentally redesign how their APIs operate.
The discovery exposes a fundamental tension in how frontier AI companies balance openness with security. By design, these models send encrypted reasoning traces to users' machines to offload computation—a practical necessity for cost and speed. Yet encryption alone, it turns out, is insufficient protection when companies also maintain a family of related models of different sizes. The attack exploits this ecosystem asymmetry: smaller models, trained with less stringent safety alignment, act as decryption keys for the larger models' proprietary reasoning.
The timing of this disclosure maps onto a wider geopolitical anxiety about model distillation. In February and June, OpenAI and Anthropic respectively alleged that Chinese companies (DeepSeek and Alibaba) had distilled their models to build competing AI systems. This research does not prove that Moonshot AI distilled from US models—the authors explicitly state the work "cannot causally establish distillation"—but the striking similarity in reasoning outputs to Claude and GPT systems on certain prompts, combined with speculation on Chinese social media that hidden reasoning traces could be used for distillation, suggests the technique may be both feasible and potentially already in use.
The policy implication cuts both ways. US policymakers view distillation as intellectual property theft that gives China a shortcut to parity. Conversely, figures like Meta CEO Mark Zuckerberg argue that distillation is a standard, beneficial practice that accelerates progress for the entire ecosystem—and that blocking it would handicap US open-source development. Researcher Kyle Miller at CSET noted uncertainty about distillation's actual strategic value, suggesting that even without the technique, Chinese labs likely possess the expertise to build models from scratch. The broader lesson: as reasoning models mature, the boundary between transparency (useful for debugging and improvement) and secrecy (necessary for competitive advantage) becomes harder to defend with encryption alone.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Amazon and Google are intensifying competitive efforts against The Trade Desk (TTD), a major digital advertisi…

OpenAI introduced Premium Seats for ChatGPT Business, priced at $125 per user per month ($100 with annual bill…

Anthropic pledged to embed machine-readable watermarks in Claude-generated text and digitally signed provenanc…

Anthropic has signed the EU AI Act Code of Practice and will embed invisible watermarks in Claude-generated te…

Anthropic has agreed to pay $9.1 billion over 20 years to Riot Platforms Inc., a Bitcoin miner turned data cen…

Cloudflare announced its AI Agents platform on August 4, introducing a two-tier wallet system—Account Wallets…

The AI news that matters, in one minute each morning.
Sign up free