
Z.ai's open-weight model GLM-5.2 has matched the cyber and biological capabilities of frontier AI systems like OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7, but unlike those closed models, it lacks effective safety guardrails.
When evaluated by SaferAI, GLM-5.2 refused none of the offensive tasks it was given, while competing frontier models consistently refused such requests.
This gap matters because open-weight models can be downloaded and run on any hardware, allowing users to strip away safety protections—a risk that intensifies as open-weight AI approaches the power of the industry's leading systems.
What happened
Z.ai's GLM-5.2, an open-weight AI model, is only a few months behind OpenAI's GPT-5.5 and Anthropic's Claude Opus 4.7 in cyber and biological capabilities, according to SaferAI's evaluation. However, GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was tested on, while Claude Opus 4.7 refused them so consistently that SaferAI could not complete the CyberGym benchmark on it.
Why it matters
Open-weight models allow anyone to download and run the AI weights on their own hardware, where they can remove or modify safeguards—a risk frontier developers like OpenAI and Anthropic try to manage through classifiers, refusal training, and API-level controls. Z.ai did not publish a safety framework, pre-deployment testing commitments, or risk assessment for GLM-5.2, raising concerns that highly capable AI could be accessed by attackers with no oversight once released.
What to watch
The debate is shifting from whether open-weight models can match frontier capabilities to how society manages the risks they pose. SaferAI's executive director noted that "the frontier of capability is not the frontier of risk," and advocates for techniques like pre-training data filtering to reduce hazardous knowledge without harming model performance—though such filtering is less practical for cybersecurity than for biology.
Ask the AI about this article →
The release of GLM-5.2 marks a turning point in the open-weight AI debate. For years, critics warned that open-weight models could put highly capable AI into the hands of attackers with no way to police its use once downloaded. SaferAI's evaluation confirms that concern is no longer theoretical: a Chinese model has closed the gap with frontier systems on dangerous capabilities—cyber and biological—while providing no safety guardrails to match. The division is stark. Frontier developers like OpenAI and Anthropic use classifiers, refusal training, and API-level controls to limit dangerous assistance. But those measures are designed for closed systems where the developer retains control. Once weights are released, those protections vanish.
The challenge for the industry is structural. For cybersecurity, filtering training data to remove offensive knowledge is impractical: coding and hacking overlap so much that a model excelling at one will be good at the other, and coding has become AI's biggest commercial driver. Developers face pressure to keep improving those capabilities even as they search for ways to limit misuse. Frontier models have responded with selective restrictions—for example, Anthropic's Opus 5 can search for vulnerabilities in source code but not compiled software—but such measures only work on closed systems. The debate is now moving from capability parity to risk management at scale: how do you ensure that good, safe capabilities are widely accessible while keeping dangerous ones out of reach when the model code itself is public?
For example, today's edition would include:
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytimeWhat is AIToday? →
Ask AI anything about this article. Q&As are published on this page for other readers too.
Israeli startup DataAgent Ltd
SK Hynix presented a custom HBM concept at SEMICON Taiwan 2026, where compute functions are placed in the base…

The U.S. Department of Defense announced on August 31 that it has deployed ChatGPT Mil, a customized version o…

Nvidia reported earnings that were both remarkable and boring, reflecting its focus on avoiding a consolidated…

Anthropic has agreed to a $35bn cloud-computing contract with Lambda, a Nvidia-backed cloud provider

The Supreme Court of Japan has included about ¥60 million in its fiscal 2027 budget request for AI-related exp…
