
DeepSeek V4 Pro discovered more vulnerabilities than any frontier model when pooled across three runs, reaching 28 of 32.
Open-weight models now match or beat closed models on recall while costing a fraction as much.
Repeating cheap models can deliver better coverage than a single expensive run, though false-positive filtering adds downstream cost.
What happened
A cybersecurity firm benchmarked 10 AI models (including DeepSeek V4 Pro, Qwen, GLM-5.3, Grok, Opus 5, and others) across 32 real-world vulnerabilities, with each model attempting the task three times. DeepSeek V4 Pro 0813 found the most vulnerabilities when results were pooled across runs—28 of 32—while GLM-5.3 matched Opus 5 at 26 of 32 found while costing 69.8% less.
Why it matters
Open-weight models (DeepSeek, Qwen, GLM, Kimi) have now matched or exceeded closed frontier models (Opus 5, Grok 4.6, Sol) in pooled vulnerability discovery—reversing the prior hierarchy. Three runs of cheaper DeepSeek Flash ($108 total) reached 24 vulnerabilities, matching Grok's single best pass that cost $450–$590. This changes the cost-benefit calculus for security teams choosing which AI model to deploy.
What to watch
The benchmark consumed 11.7 billion tokens across 96 total runs (three per model). DeepSeek Pro's individual consistency lagged (only 10 of 28 vulnerabilities found in every run), but repetition compensated; open models also produced more false leads, shifting triage burden downstream. The researchers stress that open-weight models still require careful harness engineering to replace frontier models without human oversight.
Ask the AI about this article →
The benchmark represents a significant shift in the publicly available AI model hierarchy for security tasks. Prior evaluation showed open-weight models catching up; this benchmark demonstrates they have now caught up and in some cases surpassed frontier models on the specific task of vulnerability discovery. The key insight is that vulnerability discovery—unlike many AI tasks—benefits from output variance: running the same model multiple times on the same problem yields different exploration paths, and pooling those results recovers more vulnerabilities than any single run. DeepSeek V4 Pro found only 17 vulnerabilities in its first pass but reached 28 across three runs; Qwen climbed from 19 to 26. This property flips the usual penalty of model inconsistency into an advantage.
Cost efficiency emerges as a critical new variable. Three runs of DeepSeek Flash cost $108 and recovered 24 vulnerabilities—the same as Grok's best single run at roughly 10 times the price. GLM-5.3, tested as an early evaluation partner, improved from preview to final release, matching Opus 5's recall at 69.8% lower cost. However, the trade-off is noise: open-weight models produced more false leads, which the researchers explicitly call out as a downstream burden. This means cost savings shift the human effort from running the model to filtering its output, a detail that complicates the simple "cheaper is better" narrative.
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Chinese large-model developer Z.ai says it can now support large-scale inference using roughly 100,000 domesti…

Analyst Ming-Chi Kuo says Nvidia has revived the Rubin CPX AI accelerator with a substantially redesigned arch…

A UK study by UK AI Security Institute and Limbic AI surveyed 6,474 British adults

Broadcom's Clayton Donley says companies are doing mission-critical work with AI agents quickly, but without t…
Bank of England governor Andrew Bailey warned that advanced AI poses risks to financial infrastructure in a le…
OpenAI released a new evaluation framework on July 17, 2026, urging companies to measure AI ROI by 'useful out…
