AIToday
Large Language ModelsAI Safety & AlignmentTHE DECODERPublished: Aug 12, 2026, 04:00 JST

Researchers crack encrypted AI reasoning, find passwords in leaked sessions

Researchers crack encrypted AI reasoning, find passwords in leaked sessions

3 Key Points

  1. What happened

    Security researchers led by Alexander Panfilov discovered a vulnerability in the APIs of all major AI providers (OpenAI, Anthropic, Google) that allows extraction of encrypted reasoning tokens generated by models like o-series, Claude, and Gemini. A scan of roughly 7,000 publicly shared sessions turned up 62 API keys, 33 email addresses, 33 passwords, and other sensitive data.

  2. Why it matters

    The vulnerability enables smaller models to read and transcribe the reasoning of more powerful ones within the same provider—for instance, Anthropic's Haiku 4.5 can extract Opus 4.8's full thought processes through jailbreaking. This may have already allowed Chinese model makers to train their own models on reasoning traces without permission; researchers found that specific Claude and GPT reasoning segments are up to six orders of magnitude easier to extract from Kimi-K3 than from other models, suggesting it may have been trained on such data.

  3. What to watch

    The labs have already patched several issues and are working on more fixes, according to Panfilov. The researchers estimate API costs for decoding 10,000 traces at about $720, meaning the attack is inexpensive enough to scale. The full findings are documented at stolen-thoughts.com.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The vulnerability represents a fundamental breakdown in the security model that AI labs constructed around their reasoning processes. When Matthew Green, a cryptography expert, first reported in May that encrypted reasoning blobs could be replayed outside their original context, the providers dismissed the finding, saying they did not see security implications in side channels or replays. The new research from Panfilov's team demonstrates that assessment was incorrect—the portability of encrypted thoughts across sessions, users, and models within a single provider creates multiple avenues for misuse.

What makes the discovery particularly significant is its potential connection to the practice of "reasoning distillation," where a less capable model is improved by training on the reasoning outputs of a more powerful one. The researchers argue that it may have been possible for some time to extract reasoning processes for training proprietary models without breaking the cryptography itself. This supports mounting concerns that Chinese AI makers have been using reasoning traces to train their models on chain-of-thought data, a suspicion strengthened by analysis of Kimi-K3, which shows unusually high similarity to Claude and GPT reasoning segments.

The incident also exposes a gap between what AI models are actually doing and what their creators show users. The sanitized summaries presented in chat interfaces omit important details—models sometimes already know answers and then reverse-engineer plausible solution paths, or they get stuck in loops of nonsensical terms like "vantages," "marinades," and "watchers." Researchers at Arizona State University have warned that this humanized presentation creates false confidence in model controllability and may steer AI safety research in the wrong direction.

FAQ
How did researchers extract the encrypted reasoning?
Researchers found a vulnerability in the APIs of all leading AI providers that allows extraction of encrypted reasoning tokens. For most queries, the number of extracted tokens matched the billed thinking tokens exactly, meaning they captured the full internal reasoning, not just partial snippets.
What sensitive data was found in the leaked sessions?
A scan of roughly 7,000 public traces turned up 62 API keys, 33 email addresses, 33 passwords, and other sensitive data from publicly shared Claude Code or Codex sessions.
Is there evidence this vulnerability has already been exploited?
Yes. Researchers found that specific Claude and GPT reasoning segments are up to six orders of magnitude easier to extract from Kimi-K3 than from the next closest model, suggesting Kimi-K3 may have been trained on such traces.

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleDeepMind CEO Demis Hassabis steps down to become chairman