AIToday
Large Language ModelsAI Safety & AlignmentITmedia AI+Published: Oct 1, 2026, 16:00 JST

OpenAI blocks mass reasoning extraction, ties core to Moonshot AI

OpenAI blocks mass reasoning extraction, ties core to Moonshot AI

3 Key Points

  1. What happened

    OpenAI said it blocked a coordinated effort to extract protected reasoning from its models, which it calls "adversarial distillation," and concluded the core of the activity was tied to people connected to Moonshot AI, maker of Kimi.

  2. Why it matters

    Extracted reasoning can train another model without inheriting the original's safety measures, so OpenAI says adversarial distillation carries safety and national security risks — a concern it expects to grow as model abilities spread into dual-use areas.

  3. What to watch

    Whether the activity was run by a single group is still unknown, and OpenAI says the technique is not specific to its models — so the test is whether other providers harden hosted environments the same way.

WHO IT HITSCompanies hosting or reselling frontier models on their own clouds, and the security teams responsible for watermarking and protecting hidden reasoning, may need to match OpenAI's protections. Model providers training on rivals' outputs also face tighter detection.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

FAQ
Did the attackers break OpenAI's encryption or hack its systems?
No. OpenAI says they did not break encryption, break into databases, or access stored user conversations. Instead they manipulated interactions with the model so hidden reasoning became visible to the requester.
What exactly is the new technique OpenAI found?
In one case, encrypted reasoning from one conversation was copied into another, and the model was made to decrypt and transcribe that content. OpenAI calls this a new method.
Who is "adversarial distillation" a risk to?
OpenAI says it carries safety and national security risks. Someone can train another model using extracted reasoning without inheriting the original's safety measures, and advanced capabilities can transfer without equal investment in safety.

Also reported by AI Watch (Impress)

AI news that matters for your work, delivered every morning.

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleAmazon signs 20-year nuclear deal with Constellation