AIToday
Large Language ModelsAI Safety & AlignmentML Safety NewsletterPublished: Jul 23, 2026, 01:00 JST

Frontier LLMs can build working exploits from known bugs, benchmarks show

Frontier LLMs can build working exploits from known bugs, benchmarks show

3 Key Points

  1. What happened

    Two new benchmarks—ExploitGym and ExploitBench—demonstrate that advanced AI systems can convert knowledge of software vulnerabilities into functional exploits. Mythos Preview achieved full arbitrary code execution (Tier 1) on 18 different bugs, while GPT-5.5 achieved it in only one case.

  2. Why it matters

    AI systems already discover critical vulnerabilities across major operating systems and software; the ability to automatically build exploits from those vulnerabilities closes a significant remaining barrier to widespread automated cyberattacks. Frontier AIs may soon bridge this gap, creating risk of large-scale attacks on critical infrastructure.

  3. What to watch

    ExploitGym involves vulnerabilities in Linux, V8, and other software with real-world defenses like sandboxing; ExploitBench focuses solely on V8 and measures a five-tier capability ladder from code interaction up to full machine control. Researchers note that publicly available AIs currently exploit only a small fraction of known vulnerabilities, but this may change in coming months.

Not sure about something? Ask the AI

Summaries like this, in your inbox every morning.

Context & Analysis

The emergence of these two benchmarks marks a critical juncture in AI security research. AI systems have already demonstrated the ability to discover novel vulnerabilities in widely deployed software—operating systems, browsers, and development tools. The bottleneck has been exploit construction: the technical work needed to turn a known vulnerability into an attack that actually works in real-world conditions with genuine defenses like sandboxing. ExploitGym and ExploitBench measure exactly this capability, filling a gap left by previous cybersecurity benchmarks that researchers note have become saturated.

The disparity between Mythos Preview's 18 successful Tier 1 exploits and GPT-5.5's single success indicates meaningful variation in frontier model capabilities, but both results point in the same direction. The researchers explicitly observe that publicly available AI today can only exploit a small fraction of known vulnerabilities—a constraint that the coming generation of frontier models may overcome. The implication is direct: if AI exploit-building capability reaches parity with AI vulnerability discovery, the preconditions for automated cyberattacks at scale will exist. This is not speculation about new attack methods but automation of an existing human process using tools already within reach.

FAQ
What is the difference between ExploitGym and ExploitBench?
ExploitGym, led by UC Berkeley, Max Planck Institute, and UC Santa Barbara researchers, focuses on multiple software targets (Linux kernel, V8, and others) and a single piece of information to exfiltrate. ExploitBench, developed by Carnegie Mellon University, focuses only on V8 and measures a wide range of malicious capabilities on a five-tier capability ladder from code interaction to full arbitrary code execution.
Which AI system performed best at building exploits?
Mythos Preview achieved Tier 1 (arbitrary code execution) access through exploiting 18 different bugs, whereas GPT-5.5 achieved Tier 1 access in only one case.
What does the five-tier capability ladder in ExploitBench measure?
It measures escalating malicious harm: Tier 5 covers code interaction and patch attempts; Tier 4 involves bug triggering; Tier 3 means manipulation within the V8 sandbox; Tier 2 allows control outside the sandbox; Tier 1 is full arbitrary code execution on the compromised machine.
ML Safety NewsletterRead Original Article

Get the latest Large Language Models news every morning

For example, today's edition would include:

  • Abeam and Notion target enterprise knowledge for AI agentsITmedia AI+ · 1h ago
  • Zscaler unveils Agentic SOC with AI agentsITmedia AI+ · 1h ago
  • Generative Partners launches "AX BPO" for work AI alone can't finishITmedia AI+ · 1h ago

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleKimi K3 release sparks open-model momentum; China commits to openness strategy