AIToday

UK: gap between open and closed AI models on cybersecurity shrinking

Import AI15h ago
UK: gap between open and closed AI models on cybersecurity shrinking

Key takeaway

The UK government's AI Security Institute has reported that open-weight AI models are rapidly closing the cybersecurity capability gap with proprietary ones, with recent models performing at levels that proprietary competitors reached 4–7 months earlier. This narrowing gap means cyber defenders have a limited window before frontier cybersecurity capabilities become widely accessible without the safeguards that proprietary companies maintain, raising urgent questions about how AI policy and security will adapt to broadly distributed powerful models.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    The UK government's AI Security Institute found that recent open-weight models GLM-5.2 and DeepSeek V4-Pro now perform similarly to frontier closed models released 4 to 7 months earlier, compared with a 6 to 10 month gap measured through most of 2025. On narrow cybersecurity tasks, GLM-5.2 matches Claude Opus 4.6 (released 4.3 months prior) and DeepSeek V4-Pro sits between Claude Opus 4.5 and GPT-5.

  • Why it matters

    As open-weight models catch up to proprietary ones in cybersecurity capabilities, defenders face a narrowing window to prepare before frontier cyber capabilities may become accessible without the safeguards used by proprietary companies. This shift from a controllable frontier to widely diffused, openly available AI systems will reshape both the threat landscape and policy discussions.

  • What to watch

    The UK AI Security Institute intends to test Kimi K3 (a Chinese 2.8 trillion parameter model) on the same cybersecurity evaluation basis once its weights are publicly released. The institute found the gap widens somewhat on long-horizon cyber tasks that chain multiple capabilities together, with GLM-5.2 reaching Claude Opus 4.5 level while DeepSeek V4-Pro falls below Sonnet 4.5.

In Depth

The UK government's AI Security Institute published the first public analysis of how far leading open-weight models trail proprietary frontier models specifically on cybersecurity tasks. The results show a marked acceleration in open-source capability development.

On a benchmark of 70 evals for narrow, specific cybersecurity capabilities, GLM-5.2 performed closest to Claude Opus 4.6, which was released 4.3 months earlier. DeepSeek V4-Pro positioned itself between Claude Opus 4.5 (released in August 2025) and GPT-5 (released in November 2025). This represents a compression of the historical gap: through most of 2025, the Institute measured lags of 6 to 10 months. The Institute intends to evaluate Kimi K3, a Chinese model, once its weights become publicly available.

For longer-horizon cyber range tasks—which test an AI system's ability to chain multiple capabilities together to complete full hacking operations—the gap widens somewhat. GLM-5.2 reaches the level of Opus 4.5 (released less than 7 months before), while DeepSeek V4-Pro falls below Sonnet 4.5, a model released 7 months earlier. The Institute attributes this larger gap to open models' sometimes lacking the "generalization magic juice" that distinguishes proprietary systems, a phenomenon industry participants call "big model smell."

The Institute's central concern is temporal: as the frontier-to-open lag shrinks, "cyber defenders have a short window to prepare before today's frontier cyber capabilities may become accessible without the same safeguards" employed by proprietary companies. Once open-weight models with frontier cybersecurity capabilities are released, they become permanently distributed and impossible to recall, fundamentally shifting the balance between offense and defense.

Context & Analysis

The UK AI Security Institute's analysis reveals a structural shift in the AI landscape: the technical capability gap between closely guarded proprietary models and freely available open-weight alternatives is contracting faster than previously observed. Through 2025, this gap averaged 6 to 10 months; now it has narrowed to 4 to 7 months for cutting-edge capabilities. This trend is driven partly by aggressive open-source development from Chinese firms (exemplified by DeepSeek and Kimi), but the implications extend beyond geopolitical competition.

The significance lies not in the raw performance numbers, but in the control surface. Proprietary models are deployed by a small number of companies that can theoretically gate access, apply safety classifiers, and enforce know-your-customer procedures. Open-weight models, once released, cannot be recalled or centrally controlled. As the gap narrows, widely distributed, uncontrollable AI systems with cybersecurity capabilities—one of the most sensitive domains—become inevitable. The Institute notes that long-horizon cyber tasks (which chain multiple capabilities into full hacking operations) still show a wider gap, suggesting open models may still lack some generalization ability; but this too will likely close over time.

FAQ

Which open-weight models are now closest to proprietary frontier models on cybersecurity?
GLM-5.2 is closest to Claude Opus 4.6 (released 4.3 months earlier), while DeepSeek V4-Pro sits between Claude Opus 4.5 and GPT-5. Both models now perform similarly to proprietary frontier models released 4 to 7 months before them.
What will Kimi K3 be released as?
Kimi K3, a 2.8 trillion parameter model, will be released as an open-weight model with its weights made available in the coming weeks along with a research paper about the model.
Why is the shrinking gap between open and closed models significant?
The UK AI Security Institute warns that cyber defenders have a short window to prepare before frontier cyber capabilities may become accessible without the same safeguards used by proprietary companies, fundamentally changing the balance between offense and defense.

Get AI news like this every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No discussion yet for this article

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →