
What happened
On Anthropic's ExploitBench, GLM-5.3 built a working Chrome V8 exploit in 50 of 410 attempts versus Mythos Preview's 56, and it reached full control in 4 percent of tasks on Anthropic's internal benchmark versus 6 percent.
Why it matters
Anthropic's own numbers put an openly downloadable model this close to a guarded one on exploit development, so defenders can no longer count on the frontier gap alone to slow attackers.
What to watch
The comparison still hinges on who is counted as frontier, since Anthropic's side includes models only vetted users can access and was tested with cyber safeguards off. Watch how fast unlocked GLM-5.3 variants spread.
WHO IT HITSSecurity teams and vulnerability researchers at software vendors now face open-weight models that can develop working exploits at near-frontier levels, while government testing bodies weigh whether to act on Anthropic's call to vet capable models.
Summaries like this, in your inbox every morning.
Anthropic's report lands in a debate that was already running. The UK's AI Security Institute recently found that open models had narrowed their cyber-capability lag from six to ten months down to four to seven months, and warned that their safeguards are largely ineffective. What was unclear then was whether open models would also catch up to the leap that Claude Mythos Preview represented. The ExploitBench and internal binary exploitation measurements from Anthropic, plus CAISI's independent assessment, suggest GLM-5.3 is a first answer to that question.
Anthropic frames its own release choices as the counterfactual. It deliberately held Mythos Preview back, giving access only to select defenders through Project Glasswing, who the company says have since found more than 10,000 vulnerabilities in critical software; OpenAI is taking a similar approach with Daybreak. GLM-5.3, by contrast, is available for anyone to download, and Anthropic says several developers released unlocked versions within days of its launch. Its abliteration experiment showed how quickly refusals can be stripped, though the simulation does not execute code and so cannot show whether an attack would actually have succeeded.
The report is not a neutral document. Anthropic does not release its model weights and presents that as a security advantage, and a cheap Chinese open-weight model close to the frontier is a direct competitor; its call for government testing of GLM-5.3's successors also invites suspicion of regulatory capture. Still, the capability numbers are backed by CAISI, and unlocked versions are already out there, so the practical question for defenders and governments is likely to be less whether the gap has narrowed than how fast they can prepare while it still exists.
Pick your industry and the AI tools you use, and get news related to your work every day.
Free · 30 seconds with Google · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. The AI reads this article, earlier AIToday articles, and Wikipedia, and cites its sources. Q&As are published on this page for other readers too.
Ascerta Inc. raised $18 million in a Series A led by Dell Technologies Capital, with Hitachi Ventures, BGV and…
Netlist has begun legal proceedings against Micron and downstream customers Nvidia, Broadcom, and Google over…

Anthropic filed a confidential draft prospectus with a 2025 revenue of $4.6 billion, up from $400 million in 2…

Google is paying about 100 digital publishers for content used in AI Overviews, AI Mode, and Gemini, with paym…

A Qiita walkthrough trained a five-label car-damage classifier on Gemini Enterprise Agent Platform AutoML usin…

Lauren Tan says she shipped about 2,000 pull requests a month to production on the SpaceX AI Grok Bot team
