AIToday
Large Language ModelsAI Safety & AlignmentOpen-Source AIHacker NewsPublished: Aug 21, 2026, 22:00 JST3 min read

DeepSeek V4 Pro tops 10 AI models in cybersecurity benchmark at fraction of frontier cost

DeepSeek V4 Pro tops 10 AI models in cybersecurity benchmark at fraction of frontier cost

Key takeaway

  • DeepSeek V4 Pro discovered more vulnerabilities than any frontier model when pooled across three runs, reaching 28 of 32.

  • Open-weight models now match or beat closed models on recall while costing a fraction as much.

  • Repeating cheap models can deliver better coverage than a single expensive run, though false-positive filtering adds downstream cost.

3 Key Points

  1. What happened

    A cybersecurity firm benchmarked 10 AI models (including DeepSeek V4 Pro, Qwen, GLM-5.3, Grok, Opus 5, and others) across 32 real-world vulnerabilities, with each model attempting the task three times. DeepSeek V4 Pro 0813 found the most vulnerabilities when results were pooled across runs—28 of 32—while GLM-5.3 matched Opus 5 at 26 of 32 found while costing 69.8% less.

  2. Why it matters

    Open-weight models (DeepSeek, Qwen, GLM, Kimi) have now matched or exceeded closed frontier models (Opus 5, Grok 4.6, Sol) in pooled vulnerability discovery—reversing the prior hierarchy. Three runs of cheaper DeepSeek Flash ($108 total) reached 24 vulnerabilities, matching Grok's single best pass that cost $450–$590. This changes the cost-benefit calculus for security teams choosing which AI model to deploy.

  3. What to watch

    The benchmark consumed 11.7 billion tokens across 96 total runs (three per model). DeepSeek Pro's individual consistency lagged (only 10 of 28 vulnerabilities found in every run), but repetition compensated; open models also produced more false leads, shifting triage burden downstream. The researchers stress that open-weight models still require careful harness engineering to replace frontier models without human oversight.

Ask the AI about this article →

Context & Analysis

The benchmark represents a significant shift in the publicly available AI model hierarchy for security tasks. Prior evaluation showed open-weight models catching up; this benchmark demonstrates they have now caught up and in some cases surpassed frontier models on the specific task of vulnerability discovery. The key insight is that vulnerability discovery—unlike many AI tasks—benefits from output variance: running the same model multiple times on the same problem yields different exploration paths, and pooling those results recovers more vulnerabilities than any single run. DeepSeek V4 Pro found only 17 vulnerabilities in its first pass but reached 28 across three runs; Qwen climbed from 19 to 26. This property flips the usual penalty of model inconsistency into an advantage.

Cost efficiency emerges as a critical new variable. Three runs of DeepSeek Flash cost $108 and recovered 24 vulnerabilities—the same as Grok's best single run at roughly 10 times the price. GLM-5.3, tested as an early evaluation partner, improved from preview to final release, matching Opus 5's recall at 69.8% lower cost. However, the trade-off is noise: open-weight models produced more false leads, which the researchers explicitly call out as a downstream burden. This means cost savings shift the human effort from running the model to filtering its output, a detail that complicates the simple "cheaper is better" narrative.

FAQ

How much cheaper are open-weight models than frontier models?
GLM-5.3 final release costs 69.8% less than Opus 5 while finding 25 of 32 vulnerabilities compared to Opus's 26. Three runs of DeepSeek Flash cost $108 and reach 24 vulnerabilities, matching Grok's single best pass at $450–$590.
Which model was most consistent across multiple runs?
Grok 4.6 was the most consistent closed model, finding 21 of 32 vulnerabilities in all three runs. Among open-weight models, GLM-5.3 final release found 18 of 32 consistently across all three runs.
Can you replace Opus or Sol with an open-weight model directly?
No. The researchers state open-weight models are formidable but "not yet at the calibre to replace a frontier without a carefully designed harness." Open models require more sophisticated engineering than minimal-harness frontier models.

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Nvidia revives Rubin CPX chip with major redesignYahoo Finance AI · 2h ago
  • AI advice followed by 79%, but well-being unchangedITmedia AI+ · 5h ago
  • Enterprises face agent governance gapSiliconANGLE AI · 8h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleMeta AI glasses detector apps surge as recording concerns grow