AIToday

Frontier LLMs can build working exploits from known bugs, benchmarks show

ML Safety Newsletter2h ago
Frontier LLMs can build working exploits from known bugs, benchmarks show

Key takeaway

Two new benchmarks show that frontier large language models can transform knowledge of software vulnerabilities into working exploits on a meaningful fraction of targets. Mythos Preview achieved full machine control on 18 different bugs, while GPT-5.5 achieved it once. Because AI systems already find critical vulnerabilities across major software, the ability to automate exploit construction could enable widespread cyberattacks on critical infrastructure.

Summaries like this, in your inbox every morning.

Sign up free →

3 Key Points

  • What happened

    Two new benchmarks—ExploitGym and ExploitBench—demonstrate that advanced AI systems can convert knowledge of software vulnerabilities into functional exploits. Mythos Preview achieved full arbitrary code execution (Tier 1) on 18 different bugs, while GPT-5.5 achieved it in only one case.

  • Why it matters

    AI systems already discover critical vulnerabilities across major operating systems and software; the ability to automatically build exploits from those vulnerabilities closes a significant remaining barrier to widespread automated cyberattacks. Frontier AIs may soon bridge this gap, creating risk of large-scale attacks on critical infrastructure.

  • What to watch

    ExploitGym involves vulnerabilities in Linux, V8, and other software with real-world defenses like sandboxing; ExploitBench focuses solely on V8 and measures a five-tier capability ladder from code interaction up to full machine control. Researchers note that publicly available AIs currently exploit only a small fraction of known vulnerabilities, but this may change in coming months.

In Depth

Two independent research teams have developed benchmarks that measure how well frontier large language models can convert software vulnerabilities into working exploits—a crucial capability for launching automated cyberattacks.

ExploitGym, created by researchers from UC Berkeley, the Max Planck Institute for Security and Privacy, and UC Santa Barbara, tests AI agents on known vulnerabilities from the Linux kernel, V8 (a major JavaScript engine), and other software. The agents are given access to a remote computer running vulnerable code and tasked with remotely exploiting the vulnerabilities to execute malicious code and exfiltrate secret information. The evaluation occurs against real-world security measures, including strict sandboxing around the vulnerable code. The benchmark focuses on multiple software targets but a single exfiltration objective.

ExploitBench, developed by Carnegie Mellon University, takes a narrower but deeper approach. It provides AI agents access to an out-of-date system running a sandboxed, buggy version of V8, along with a patch script for the bug. The benchmark measures success across a "capability ladder" of five tiers, ranked by the harm an attacker can cause. Tier 5 (the lowest) measures whether the LLM agent even interacts with buggy code or attempts to apply the patch. Tier 4 measures bug triggering. Tier 3 involves manipulating objects within the V8 sandbox. Tier 2 allows control outside the sandbox. Tier 1 (the highest) represents full arbitrary code execution on the compromised machine. On this ladder, Mythos Preview achieved Tier 1 access through exploiting 18 different bugs, while GPT-5.5 achieved Tier 1 in only one case.

The researchers' key finding is that AI systems already capable of discovering critical vulnerabilities now face a narrowing window before they can also automate exploit construction. Currently, publicly available AI systems can only exploit a small fraction of known vulnerabilities, limiting real-world harm. However, the researchers predict that frontier AIs in the coming months may close this gap, enabling widespread cyberattacks and potentially compromising critical infrastructure. They note that previous generations of cybersecurity benchmarks have become saturated, making exploit-focused benchmarks like ExploitGym and ExploitBench essential for tracking the evolution of frontier models' offensive capabilities.

Context & Analysis

The emergence of these two benchmarks marks a critical juncture in AI security research. AI systems have already demonstrated the ability to discover novel vulnerabilities in widely deployed software—operating systems, browsers, and development tools. The bottleneck has been exploit construction: the technical work needed to turn a known vulnerability into an attack that actually works in real-world conditions with genuine defenses like sandboxing. ExploitGym and ExploitBench measure exactly this capability, filling a gap left by previous cybersecurity benchmarks that researchers note have become saturated.

The disparity between Mythos Preview's 18 successful Tier 1 exploits and GPT-5.5's single success indicates meaningful variation in frontier model capabilities, but both results point in the same direction. The researchers explicitly observe that publicly available AI today can only exploit a small fraction of known vulnerabilities—a constraint that the coming generation of frontier models may overcome. The implication is direct: if AI exploit-building capability reaches parity with AI vulnerability discovery, the preconditions for automated cyberattacks at scale will exist. This is not speculation about new attack methods but automation of an existing human process using tools already within reach.

FAQ

What is the difference between ExploitGym and ExploitBench?
ExploitGym, led by UC Berkeley, Max Planck Institute, and UC Santa Barbara researchers, focuses on multiple software targets (Linux kernel, V8, and others) and a single piece of information to exfiltrate. ExploitBench, developed by Carnegie Mellon University, focuses only on V8 and measures a wide range of malicious capabilities on a five-tier capability ladder from code interaction to full arbitrary code execution.
Which AI system performed best at building exploits?
Mythos Preview achieved Tier 1 (arbitrary code execution) access through exploiting 18 different bugs, whereas GPT-5.5 achieved Tier 1 access in only one case.
What does the five-tier capability ladder in ExploitBench measure?
It measures escalating malicious harm: Tier 5 covers code interaction and patch attempts; Tier 4 involves bug triggering; Tier 3 means manipulation within the V8 sandbox; Tier 2 allows control outside the sandbox; Tier 1 is full arbitrary code execution on the compromised machine.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Discussion

No comments yet. Be the first to share your thoughts!

Log in to join the discussion

Related Articles

Stay ahead with AI news

Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.

Get Started Free

Free · takes 30 seconds · unsubscribe anytime

1 minute a day. The AI essentials.

200+ sources · Email / LINE / Slack

Get it free →