AIToday
AI Business & IndustryTHE DECODERPublished: Jul 24, 2026, 19:01 JST

Kimi K3 lags US models on hacking tasks; distillation may explain gap

Kimi K3 lags US models on hacking tasks; distillation may explain gap

3 Key Points

  1. What happened

    UK and US AI security institutes tested Moonshot AI's Kimi K3 on two cyber-attack benchmarks. On ExploitBench (41 Chrome V8 vulnerabilities), Kimi K3 scored 32.2 percent versus 76.2 percent for leading U.S. models and 24.4 percent for China's GLM-5.2. On a simulated 32-step network attack test called "The Last Ones," Kimi K3 reached step 17 on average, compared with 28.5 steps for U.S. leaders.

  2. Why it matters

    Kimi K3's safeguards did not block offensive cyber operations, meaning the model will assist with exploit development without resistance—a genuine security risk. However, the performance gap between Kimi K3 and leading U.S. models may stem from how the model was trained: if Kimi K3 was built partly by distilling outputs from Anthropic's Claude (whose safety classifiers block advanced cyber queries), it would miss the deeper hacking capabilities embedded in U.S. frontier models.

  3. What to watch

    Chinese open-weight models have gained cyber capabilities since early 2025 but remain 4–7 months behind leading U.S. systems, according to CAISI's Elo-based analysis. The British institute warns that growing cyber abilities in open models create "a persistent and irreversible risk of misuse," even as the gap persists.

Ask the AI about this article →

Summaries like this, in your inbox every morning.

Context & Analysis

The evaluation reveals a significant performance gap between China's Kimi K3 and leading U.S. frontier models on offensive cyber tasks, even though Kimi K3 reportedly performs competitively on standard benchmarks. This asymmetry hints at how training methodology shapes model capabilities in ways that general performance metrics do not capture. The institutes' explanation—that Kimi K3 may have been built by distilling Claude outputs, which exclude advanced offensive cyber queries due to Anthropic's safety classifiers—suggests a structural reason for the gap rather than a simple capacity shortfall. U.S. frontier models were tested with their system-level safeguards disabled to measure maximum capabilities, exposing exploit skills that are normally inaccessible through public interfaces and therefore unavailable for distillation. This disparity matters because it shows how safety architectures in training data can inadvertently limit the transfer of certain capabilities during distillation.

Chinese open-weight models have been gaining cyber capabilities since early 2025, but according to CAISI's Elo-based time-series analysis, they remain 4–7 months behind leading U.S. systems—an improvement from the 6–10 month gap recorded at the start of 2025. The institutes note that this closing gap, combined with the fact that Kimi K3's safeguards do not prevent offensive cyber assistance, creates "a persistent and irreversible risk of misuse," particularly as open-weight models become more capable.

FAQ
How much did Kimi K3 lag behind leading U.S. models on exploit development?
On the ExploitBench benchmark, Kimi K3 scored 32.2 percent, while leading U.S. models averaged 76.2 percent. On the simulated network-attack test called "The Last Ones," Kimi K3 reached step 17 out of 32 on average, compared with 28.5 steps for leading U.S. models.
Why might Kimi K3 perform worse at cyber tasks if it matches U.S. models on standard benchmarks?
Researchers suggest Kimi K3 may have been trained partly by distilling Anthropic's Claude outputs for general knowledge and programming tasks. Because Claude's safety classifiers specifically block advanced offensive cyber queries, those capabilities would be underrepresented in a distillation dataset, allowing Kimi K3 to match Western models on standard benchmarks while lacking their deeper exploit skills.
Did Kimi K3's safeguards block offensive cyber tasks?
No. Kimi K3's safeguards did not block exploit development or offensive cyber operations, and the model assisted with both without pushback.

Get the latest AI Business & Industry news every morning

For example, today's edition would include:

  • Aron launches with $8 million to automate procurement RFQsSiliconANGLE AI · 39m ago
  • Talkdesk's Andrade: 98% deploy AI, only 15% orchestrateSiliconANGLE AI · 39m ago
  • Kerravala: resolution quality is contact center AI's new unit of valueSiliconANGLE AI · 39m ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleSLB poised to cash in on Middle East oil recovery and AI data center boom