AIToday
AI Business & IndustrySimon Willison's WeblogPublished: Sep 30, 2026, 10:01 JST

Anthropic: GLM-5.3 hits 4% cyber-exploit rate, a new threshold

Anthropic: GLM-5.3 hits 4% cyber-exploit rate, a new threshold

3 Key Points

  1. What happened

    Anthropic's Frontier Red Team evaluated several models on 100 randomly selected tasks from its internal Binary Exploitation benchmark and found GLM-5.3 achieved full control flow hijacks in 4% of trials, versus 6% for Claude Mythos Preview.

  2. Why it matters

    Reaching a full control flow hijack at all is new here — the team says a meaningful threshold has clearly been crossed because earlier models showed zero successes on these same tasks.

  3. What to watch

    The evaluation covers only 100 randomly selected tasks from an internal benchmark, so how broadly these success rates generalize is an open question; the team itself frames the result as one threshold, not a full capability picture.

WHO IT HITSSecurity teams tracking AI-enabled offensive tooling should note that a leading open-weights model now completes real exploitation tasks in some trials, even while trailing Anthropic's own frontier model. The shift is that the earlier models scored zero, so the capability floor has moved.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

The comparison the Frontier Red Team draws is not between GLM-5.3 and Claude Mythos Preview in isolation, but between this generation and the previous one. When Claude Opus 4.6 and GLM-5.2 score zero on the same 100-task sample, the appearance of any successful full control flow hijack marks a step change rather than a marginal gain. The team's word choice — a meaningful threshold — signals that it treats the existence of these successes, not their size, as the headline.

The benchmark itself is internal and the task sample is small and randomly drawn, so the 4% and 6% figures are best read as evidence of capability presence rather than as precise capability rankings. Anthropic's framing acknowledges this by noting GLM-5.3 performs below Claude Mythos Preview while still describing the overall result as a crossed threshold.

The stakes, from this body of evidence, hinge on whether the successes hold up across a wider sample of tasks. For those tracking AI and cyber risk, the thing to weigh is not which model leads but that the floor has risen from zero — and the Frontier Red Team has put its own name behind that reading.

FAQ
What is the Binary Exploitation benchmark Anthropic used?
It is Anthropic's internal benchmark, and the Frontier Red Team drew 100 tasks from it at random for this evaluation.
Which model performed better on this test?
Claude Mythos Preview did, at 6% versus 4% for GLM-5.3 — but the team emphasizes that both cleared a bar earlier models did not.
Why does the Frontier Red Team call this a threshold being crossed?
Because earlier models like Claude Opus 4.6 and GLM-5.2 succeeded in none of the trials, making any nonzero success rate a qualitative change rather than a small improvement.
Simon Willison's WeblogRead Original Article

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articleOpenAI unveils Dots, a 24-hour agent on GPT-6 Astra