AIToday
Large Language ModelsAI Safety & AlignmentOpen-Source AIGIGAZINE AIPublished: Sep 30, 2026, 16:00 JST

Anthropic: open-weight GLM-5.3 nears Claude Mythos Preview on exploits

Anthropic: open-weight GLM-5.3 nears Claude Mythos Preview on exploits

3 Key Points

  1. What happened

    Anthropic published a September 29, 2026 review finding Z.ai's downloadable GLM-5.3 built working exploits in 50 of 410 ExploitBench trials (about 12%), versus Claude Mythos Preview's 56 of 410 (14%).

  2. Why it matters

    Anthropic says its own comparably capable Claude Mythos Preview was withheld from public release and limited to trusted defenders, so GLM-5.3 appears to put similar cyber capability in anyone's hands.

  3. What to watch

    The test was a simulation with no outside system access, which Anthropic itself says does not fully reproduce real-world behavior, so the practical risk hinges on how such models are used.

WHO IT HITSSecurity teams and vulnerability researchers now face a downloadable model that Anthropic says can find and weaponize flaws autonomously, while the same capability is also available to defenders patching software.

Not sure about something? Ask the AI

Questions and answers are published on this page.

Summaries like this, in your inbox every morning.

Context & Analysis

Anthropic's review rests on a comparison the company has been drawing for a while: Claude Mythos Preview, which Anthropic describes as the first AI model able to autonomously build advanced end-to-end cyber exploits, was never released publicly. Instead it was restricted to trusted cyber defenders and used through programs like Project Glasswing to find vulnerabilities in important software. GLM-5.3, by contrast, is an open-weight model anyone can download.

The two benchmarks in the review put the gap between them in narrow terms. ExploitBench, which tests whether a model can use a known vulnerability in Google Chrome's V8 JavaScript engine, produced working exploits in roughly 12% of GLM-5.3's attempts against 14% for Claude Mythos Preview. Anthropic's own benchmark of open-source software used by Google's OSS-Fuzz found control-flow-hijack success rates of 4% and 6%. Kimi K3, DeepSeek-V4.1-Flash, and the older Claude Opus 4.6 and GLM-5.2 all scored 0%.

Anthropic also reports that a smaller variant, GLM-5.3-Flash, chained two known vulnerabilities — including the public Chrome flaw CVE-2026-11645 — into an attack that bypassed pointer authentication on ARM64, with researchers spending about 20 minutes and the model running for roughly 8 hours. Separately, Anthropic says GLM-5.3's built-in refusal mechanism was relatively easy to circumvent, and that Claude models with safeguards active did not carry out harmful operations in the same simulation.

The stakes appear to hinge on how much weight to give a controlled test. Anthropic acknowledges the trials ran in a simulated environment with no connection to outside systems and do not fully reproduce real-world behavior. Its own conclusion is nonetheless a policy one: that governments should thoroughly safety-test high-performance AI models — while noting the same capabilities can serve defenders finding and fixing flaws.

FAQ
How does GLM-5.3 compare with Claude Mythos Preview on exploit building?
On ExploitBench, GLM-5.3 succeeded in 50 of 410 trials (about 12%) and Claude Mythos Preview in 56 of 410 (14%). On Anthropic's separate control-flow-hijack benchmark, GLM-5.3 scored 4% and Claude Mythos Preview 6%.
How easily can GLM-5.3's safety refusals be removed?
Anthropic says applying "abliteration" to weaken the model's internal refusal behavior cut refusal rates above 90% on JailbreakBench and HarmBench down to roughly 3% and 2%, and to about 12% on StrongREJECT, with no change in GPQA-Diamond science scores.
What did the US government's CAISI say about GLM-5.3?
Anthropic notes that CAISI, part of the National Institute of Standards and Technology, reviewed GLM-5.3 on September 17, 2026 and called it the most cyber-capable open-weight model published to date.

AI news that matters for your work, in one minute a day

Pick your industry and the AI tools you use, and get news related to your work every day.

Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →

Ask AI

Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.

Questions and answers are published on this page.

Related Articles

Next articlePalantir Technologies $186.97 share broadly matches DCF value