
What happened
Anthropic's Frontier Red Team evaluated several models on 100 randomly selected tasks from its internal Binary Exploitation benchmark and found GLM-5.3 achieved full control flow hijacks in 4% of trials, versus 6% for Claude Mythos Preview.
Why it matters
Reaching a full control flow hijack at all is new here — the team says a meaningful threshold has clearly been crossed because earlier models showed zero successes on these same tasks.
What to watch
The evaluation covers only 100 randomly selected tasks from an internal benchmark, so how broadly these success rates generalize is an open question; the team itself frames the result as one threshold, not a full capability picture.
WHO IT HITSSecurity teams tracking AI-enabled offensive tooling should note that a leading open-weights model now completes real exploitation tasks in some trials, even while trailing Anthropic's own frontier model. The shift is that the earlier models scored zero, so the capability floor has moved.
Summaries like this, in your inbox every morning.
The comparison the Frontier Red Team draws is not between GLM-5.3 and Claude Mythos Preview in isolation, but between this generation and the previous one. When Claude Opus 4.6 and GLM-5.2 score zero on the same 100-task sample, the appearance of any successful full control flow hijack marks a step change rather than a marginal gain. The team's word choice — a meaningful threshold — signals that it treats the existence of these successes, not their size, as the headline.
The benchmark itself is internal and the task sample is small and randomly drawn, so the 4% and 6% figures are best read as evidence of capability presence rather than as precise capability rankings. Anthropic's framing acknowledges this by noting GLM-5.3 performs below Claude Mythos Preview while still describing the overall result as a crossed threshold.
The stakes, from this body of evidence, hinge on whether the successes hold up across a wider sample of tasks. For those tracking AI and cyber risk, the thing to weigh is not which model leads but that the floor has risen from zero — and the Frontier Red Team has put its own name behind that reading.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
HBM is approaching half the cost of a GPU-HBM CoWoS package, prompting the question of whether memory remains…

DeepSeek is bringing more of the software it uses to develop its AI models to Huawei Technologies' Ascend 950…

Among respondents at companies with 1,001+ employees, 50.0% said AI is used company-wide, and 46.0% flagged AI…

Oracle invoked "force majeure" to delay payment on its Project Jupiter data center, and its 2056 bonds then tr…

Nvidia released the Open Agent Safety Platform on September 28, days after CEO Jensen Huang called warnings fr…

McDonald’s is increasingly using AI to guide menu prices in the U.S
