AIToday

Kimi K3 lags US models on hacking tasks; distillation may explain gap

THE DECODER4h ago
Kimi K3 lags US models on hacking tasks; distillation may explain gap

Key takeaway

UK and U.S. AI security institutes evaluated Moonshot AI's Kimi K3 on offensive cyber tasks and found it trails leading American frontier models by a wide margin—scoring 32.2 percent on exploit-development benchmarks versus 76.2 percent for U.S. leaders—but outperforms China's GLM-5.2. Kimi K3's safeguards did not block offensive cyber operations. Researchers suggest the performance gap may reflect the model's training approach: if built partly by distilling Claude's outputs, Kimi K3 would inherit Claude's general knowledge and programming skills without the deeper exploit capabilities hidden behind Claude's safety classifiers.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    UK and US AI security institutes tested Moonshot AI's Kimi K3 on two cyber-attack benchmarks. On ExploitBench (41 Chrome V8 vulnerabilities), Kimi K3 scored 32.2 percent versus 76.2 percent for leading U.S. models and 24.4 percent for China's GLM-5.2. On a simulated 32-step network attack test called "The Last Ones," Kimi K3 reached step 17 on average, compared with 28.5 steps for U.S. leaders.

  • Why it matters

    Kimi K3's safeguards did not block offensive cyber operations, meaning the model will assist with exploit development without resistance—a genuine security risk. However, the performance gap between Kimi K3 and leading U.S. models may stem from how the model was trained: if Kimi K3 was built partly by distilling outputs from Anthropic's Claude (whose safety classifiers block advanced cyber queries), it would miss the deeper hacking capabilities embedded in U.S. frontier models.

  • What to watch

    Chinese open-weight models have gained cyber capabilities since early 2025 but remain 4–7 months behind leading U.S. systems, according to CAISI's Elo-based analysis. The British institute warns that growing cyber abilities in open models create "a persistent and irreversible risk of misuse," even as the gap persists.

In Depth

The British AI Security Institute and the U.S. Center for AI Standards and Innovation jointly tested Moonshot AI's Kimi K3 on two benchmark suites designed to measure offensive cyber capabilities. The first, ExploitBench, was developed by Carnegie Mellon University and includes 41 vulnerabilities discovered in Chrome's V8 engine after 2023. The benchmark tracks how far a model advances through the software exploitation process. Kimi K3 scored 32.2 percent on this task, well below the 76.2 percent average achieved by leading U.S. models and above the 24.4 percent score of China's GLM-5.2. Notably, Kimi K3 did not reach the highest exploit level—Arbitrary Code Execution (ACE)—on any of the 41 tasks, whereas leading U.S. models achieved ACE in 20 of the 41 cases. ACE is the most severe exploit level because it grants attackers full control over a target system.

The second test, called "The Last Ones," simulated a multi-stage corporate network attack across four subnets with roughly 20 hosts and a 32-step attack path that a human expert would require approximately 20 hours to complete. Kimi K3 reached step 17 on average, compared with 28.5 steps for leading U.S. models and 11 steps for GLM-5.2. In one of ten attempts, Kimi K3 completed the entire attack path while staying within the 100 million token limit, demonstrating latent capability but unreliable performance. The institutes concluded that "Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access." Neither Kimi K3 nor GLM-5.2 achieved full exploits in this test.

Crucially, Kimi K3's safeguards did not block either exploit development or offensive cyber operations—the model assisted with both without resistance. The institutes tested leading U.S. models with their system-level safeguards disabled to measure maximum capabilities; those safeguards are enabled in publicly available versions. A parallel time-series analysis by CAISI tracks cyber capabilities of U.S. and Chinese models since early 2025 using an Elo-based scale. Both trend lines are climbing, but Chinese models consistently remain behind their U.S. counterparts. In a previous analysis, the British institute reported that the performance gap for open models had narrowed to 4–7 months compared with 6–10 months at the start of 2025.

Researchers note that Kimi K3's strong performance on general benchmarks but weak cyber scores may reflect its training approach. U.S. science advisor Michael Kratsios recently alleged that Moonshot AI had distilled Anthropic's Claude model by using Claude's outputs as training data. One explanation for the cyber gap is that Kimi K3 was trained largely on Claude outputs covering general knowledge, programming, and agent tasks—but since Anthropic's safety classifiers specifically block advanced offensive cyber queries, those capabilities would be underrepresented in the distillation dataset. This would allow Kimi K3 to match leading Western models on standard benchmarks while lacking their deeper exploit skills. The institutes' findings support this reading: by disabling system-level safeguards on U.S. models, they revealed cyber capabilities nearly impossible to access through public interfaces and therefore largely unavailable for distillation. The British institute warns that the growing cyber capabilities of open models, regardless of the gap with U.S. systems, create "a persistent and irreversible risk of misuse."

Context & Analysis

The evaluation reveals a significant performance gap between China's Kimi K3 and leading U.S. frontier models on offensive cyber tasks, even though Kimi K3 reportedly performs competitively on standard benchmarks. This asymmetry hints at how training methodology shapes model capabilities in ways that general performance metrics do not capture. The institutes' explanation—that Kimi K3 may have been built by distilling Claude outputs, which exclude advanced offensive cyber queries due to Anthropic's safety classifiers—suggests a structural reason for the gap rather than a simple capacity shortfall. U.S. frontier models were tested with their system-level safeguards disabled to measure maximum capabilities, exposing exploit skills that are normally inaccessible through public interfaces and therefore unavailable for distillation. This disparity matters because it shows how safety architectures in training data can inadvertently limit the transfer of certain capabilities during distillation.

Chinese open-weight models have been gaining cyber capabilities since early 2025, but according to CAISI's Elo-based time-series analysis, they remain 4–7 months behind leading U.S. systems—an improvement from the 6–10 month gap recorded at the start of 2025. The institutes note that this closing gap, combined with the fact that Kimi K3's safeguards do not prevent offensive cyber assistance, creates "a persistent and irreversible risk of misuse," particularly as open-weight models become more capable.

FAQ

How much did Kimi K3 lag behind leading U.S. models on exploit development?
On the ExploitBench benchmark, Kimi K3 scored 32.2 percent, while leading U.S. models averaged 76.2 percent. On the simulated network-attack test called "The Last Ones," Kimi K3 reached step 17 out of 32 on average, compared with 28.5 steps for leading U.S. models.
Why might Kimi K3 perform worse at cyber tasks if it matches U.S. models on standard benchmarks?
Researchers suggest Kimi K3 may have been trained partly by distilling Anthropic's Claude outputs for general knowledge and programming tasks. Because Claude's safety classifiers specifically block advanced offensive cyber queries, those capabilities would be underrepresented in a distillation dataset, allowing Kimi K3 to match Western models on standard benchmarks while lacking their deeper exploit skills.
Did Kimi K3's safeguards block offensive cyber tasks?
No. Kimi K3's safeguards did not block exploit development or offensive cyber operations, and the model assisted with both without pushback.

Get the latest AI Business & Industry news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime