
What happened
UK and US AI security institutes tested Moonshot AI's Kimi K3 on two cyber-attack benchmarks. On ExploitBench (41 Chrome V8 vulnerabilities), Kimi K3 scored 32.2 percent versus 76.2 percent for leading U.S. models and 24.4 percent for China's GLM-5.2. On a simulated 32-step network attack test called "The Last Ones," Kimi K3 reached step 17 on average, compared with 28.5 steps for U.S. leaders.
Why it matters
Kimi K3's safeguards did not block offensive cyber operations, meaning the model will assist with exploit development without resistance—a genuine security risk. However, the performance gap between Kimi K3 and leading U.S. models may stem from how the model was trained: if Kimi K3 was built partly by distilling outputs from Anthropic's Claude (whose safety classifiers block advanced cyber queries), it would miss the deeper hacking capabilities embedded in U.S. frontier models.
What to watch
Chinese open-weight models have gained cyber capabilities since early 2025 but remain 4–7 months behind leading U.S. systems, according to CAISI's Elo-based analysis. The British institute warns that growing cyber abilities in open models create "a persistent and irreversible risk of misuse," even as the gap persists.
Ask the AI about this article →
Summaries like this, in your inbox every morning.
The evaluation reveals a significant performance gap between China's Kimi K3 and leading U.S. frontier models on offensive cyber tasks, even though Kimi K3 reportedly performs competitively on standard benchmarks. This asymmetry hints at how training methodology shapes model capabilities in ways that general performance metrics do not capture. The institutes' explanation—that Kimi K3 may have been built by distilling Claude outputs, which exclude advanced offensive cyber queries due to Anthropic's safety classifiers—suggests a structural reason for the gap rather than a simple capacity shortfall. U.S. frontier models were tested with their system-level safeguards disabled to measure maximum capabilities, exposing exploit skills that are normally inaccessible through public interfaces and therefore unavailable for distillation. This disparity matters because it shows how safety architectures in training data can inadvertently limit the transfer of certain capabilities during distillation.
Chinese open-weight models have been gaining cyber capabilities since early 2025, but according to CAISI's Elo-based time-series analysis, they remain 4–7 months behind leading U.S. systems—an improvement from the 6–10 month gap recorded at the start of 2025. The institutes note that this closing gap, combined with the fact that Kimi K3's safeguards do not prevent offensive cyber assistance, creates "a persistent and irreversible risk of misuse," particularly as open-weight models become more capable.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
OpenAI reportedly agreed to buy smartphone camera software startup Glass Imaging for more than $300 million, p…
Aron Inc. launched with $8 million raised across two rounds
Pedro Andrade, Talkdesk's VP of AI and generative AI business specialist, told theCUBE that CXA is "an operati…
At the AI ROI in Contact Center Summit, analysts Bob Laliberte and Zeus Kerravala said contact center AI ROI w…
Nvidia invested US$3.5 billion in MediaTek convertible bonds, expanding cooperation into custom AI chips, AI P…

TechTouch surveyed 319 people overseeing generative AI at firms with over 1,000 employees
