
Researchers have released an AI Security Leaderboard that benchmarks how well frontier AI models resist jailbreak attacks — a new tool addressing the lack of comparable security rankings.
The leaderboard tests models against 1500 automatically generated jailbreak attempts and reveals substantial differences in robustness, reflecting growing concern among policymakers and developers about adversarial vulnerabilities in AI deployment.
What happened
Researchers developed an automated leaderboard that ranks AI models by security, testing them against 1500 automatically generated jailbreak attempts to measure resistance to universal jailbreaks — prompts that elicit harmful responses to over 75% of clearly harmful questions within a domain.
Why it matters
Model security is becoming critical to deployment decisions; the U.S. government has required developers to pull models for cybersecurity jailbreaks, and developers are holding back on AI agent deployments due to risks of adversarial attacks. The leaderboard found a significant gap between the most and least robust models.
What to watch
This is version 1.0; the team is soliciting feedback on methodology and considering adding open-weight models to the rankings.
Researchers have unveiled an AI Security Leaderboard designed to rank frontier models by their robustness against jailbreak attacks. Unlike the abundant capability rankings that measure how well models perform on standard benchmarks, no comparable tool existed for evaluating model security — a gap the team sought to address. The leaderboard relies on an automated test suite that runs models through 1500 automatically generated jailbreak attempts. The core metric is the count of universal jailbreaks: prompts that elicit compliant, detailed responses to more than 75% of clearly harmful questions within a specific domain (such as offensive cybersecurity). According to the team's technical report, the testing revealed a significant gap between the most and least robust models. The timing of this tool reflects broader shifts in AI governance and risk management. The U.S. government has required developers to pull models for cybersecurity jailbreaks, and developers are increasingly reluctant to deploy AI agents due to concerns about adversarial attacks. This leaderboard is version 1.0, and the researchers are actively seeking feedback from the community on methodology and future directions. One area under consideration is adding open-weight models to the rankings, though the team notes the challenge of fairly comparing them to proprietary models.
The emergence of this security leaderboard reflects a shift in how the AI industry evaluates model readiness for deployment. While capability rankings — measuring performance on standard benchmarks — have long dominated the landscape, security robustness has lagged behind as a measured attribute. The body describes a concrete gap between the most and least robust models, suggesting that blanket assumptions about frontier model safety are misplaced. The researchers' decision to test against 1500 automatically generated jailbreak attempts and focus on universal jailbreaks (those that work across over 75% of harmful prompts within a domain) establishes a quantifiable standard for what "robustness" means in practice. This matters because, as the body notes, both government actors and private developers are now making deployment decisions based on security risk — a departure from the earlier era when capability alone drove adoption.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Anthropic disclosed that three Claude models—Opus 4.7, Mythos 5, and an internal research test model—gained un…

Amazon is positioning itself as the platform provider for AI rather than competing to build the best AI model

Meta's stock fell 8% on Thursday following disappointing earnings and revenue forecasts, with the company warn…

A real-world trace of Claude Code use across a 45-person team over 30 days found that only 14% of input tokens…

Apple CEO Tim Cook said during an earnings call on Thursday that the company will offer "upgrade possibilities…

Reddit reported Q2 revenue of $805 million (up 61% year-over-year) and net income of $253 million (up 183%), b…

The AI news that matters, in one minute each morning.
Sign up free