
Google and a team of security researchers released ExploitGym, a large-scale benchmark with 869 real-world vulnerability instances designed to evaluate whether AI agents can develop exploits against userspace programs, V8, and the Linux kernel. The benchmark is open-source and actively maintained, providing the security research community with a standardized way to measure AI exploit-generation capabilities.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Google and collaborators released ExploitGym, a benchmark containing 869 instances built from real-world vulnerabilities in userspace programs, Google's V8 engine, and the Linux kernel, designed to test whether AI agents can develop working exploits.
Why it matters
The benchmark provides a systematic way to measure AI capabilities in security research—a critical area as AI agents become more autonomous. The real-world vulnerability base means results reflect practical attack scenarios rather than synthetic problems.
What to watch
The benchmark is actively maintained at v1.0; the source code is available under Apache-2.0 license, and a leaderboard submission format is documented for researchers to benchmark their AI agents.
ExploitGym is a benchmark released by Google and collaborators to evaluate AI agents' ability to develop exploits from real-world vulnerabilities. The benchmark contains 869 instances derived from actual vulnerabilities found in userspace programs, Google's V8 JavaScript engine, and the Linux kernel.
The benchmark is structured for reproducibility and ease of use. Researchers can get started by installing Python dependencies, building runtime artifacts (GDB, socat, nc, Node.js, and agent CLIs), and extracting task data via provided scripts. The setup includes pulling Docker images for target tasks, starting a controller and LLM proxy (with support for OpenAI and Anthropic APIs), and then running an agent against the benchmarked vulnerabilities. Documentation covers system dependencies, Docker image management, evaluation procedures, network isolation (via a Firewall component based on Ubuntu Squid), and submission requirements for the leaderboard.
The benchmark is actively maintained; the current release is v1.0, with a full version history available in CHANGELOG.md and the canonical task list in data/task_ids/v1.txt. The source code is released under the Apache-2.0 license, while the bundled task data retains the licenses of the external vulnerabilities it draws from. For citation, the benchmark is attributed to Wang, Schiller, Li, Sesha Narayana, Nasr, Carlini, Qi, Wallace, Bursztein, Invernizzi, Thomas, Shoshitaishvili, Guo, He, Holz, and Song, with the research published on arXiv as preprint 2605.11086.
ExploitGym represents a systematic effort to measure AI agent capabilities in a high-stakes domain: security exploit development. By anchoring the benchmark to real-world vulnerabilities rather than synthetic test cases, the benchmark grounds evaluation in practical threat scenarios. The inclusion of vulnerabilities across multiple layers—userspace programs, a major JavaScript engine, and the operating system kernel—ensures breadth of attack surface, making the benchmark a credible proxy for real-world exploit potential.
The maintenance and update cadence (currently v1.0 with 869 instances) indicates this is positioned as a long-term research infrastructure, not a one-off release. The provision of leaderboard submission infrastructure signals intent to establish a community benchmark similar to other open AI evaluation platforms. The Apache-2.0 licensing and documented setup pipeline lower barriers to adoption, though the complexity of the setup (Docker, GDB, network isolation) targets a specialized audience of security researchers and AI practitioners.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No discussion yet for this article
Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime