
Z.ai, a Beijing-based lab, released GLM 5.2 on 16 June, an open-weights AI coding model priced at $4.40 per million output tokens—roughly one-tenth the cost of Anthropic's premium models. The model performs nearly as well as Anthropic's Opus 4.8 on some benchmarks and allows companies to self-host on their own hardware, addressing both cost and data-sovereignty concerns. Early feedback shows it handles long-duration tasks and front-end work well, though some engineers report rate limits and hallucinations on complex fixes, suggesting it may be most valuable as a supplementary tool rather than a frontier-model replacement.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Z.ai released GLM 5.2, an open-weights AI coding model with 753 billion parameters (40 billion active), on 16 June. Its API costs $4.40 per million output tokens—less than a fifth of Anthropic's Opus 4.8 and a tenth of Anthropic's Fable coding model. The model nearly ties Opus 4.8 on some agentic coding benchmarks like FrontierSWE and PostTrainBench, though it lags significantly on harder tasks (scoring just 13 percent on SWE-Marathon versus Opus 4.8's 26 percent).
Why it matters
Software engineers often default to the most powerful (and expensive) models because many companies lack token budgets and cost oversight. GLM 5.2's low price and open-weights design—allowing any organization to download and host it for free on their own hardware—offers a concrete way to cut spending. For companies handling sensitive data, self-hosting avoids routing information through Chinese-linked infrastructure, a concern that has limited adoption of other Chinese models. The cost gap may test whether U.S. frontier labs' main competitive moat is technical capability or simply loose corporate spending habits.
What to watch
Real-world usage is mixed. Engineers report GLM 5.2 excels at long-horizon tasks and front-end development (handling 10–20 percent of daily work for some), but others hit rate limits within days or experience hallucinations on minor fixes. Availability depends on the access path: Z.ai's free open-weights model is available now; paid API access starts at $4.40 per million output tokens.
On 16 June, Z.ai, a Beijing-based AI lab, released GLM 5.2, an open-weights large language model with 753 billion total parameters, though only 40 billion are active at any given time—an optimization that improves response speed. The model was released under an MIT open-source license, allowing any organization to freely download, copy, modify, and distribute it. For those preferring managed access, Z.ai charges $4.40 per million output tokens via API—less than a fifth of Anthropic's Opus 4.8 pricing and a tenth of Anthropic's Fable coding model.
GLM 5.2's performance on benchmarks drew significant attention. Z.ai published a research report on launch day, comparing the model chiefly against Anthropic's Opus 4.8 and OpenAI's GPT-5.5. The model nearly matched Opus 4.8 on some agentic coding benchmarks, including FrontierSWE and PostTrainBench, and scored well on cybersecurity benchmarks, prompting comparisons to Anthropic's Mythos. However, on harder tasks, the gaps widened: GLM 5.2 completed just 13 percent of tasks in SWE-Marathon, a long-duration agentic benchmark, while Opus 4.8 achieved 26 percent. Opus 4.8 also posted 10-percent-or-greater margins on the coding benchmarks NL2Repo, DeepSWE, and Tool-Decathlon. The release sparked concern among some U.S. observers, in part because GLM 5.2 set new benchmark records for both open-weights models and Chinese-developed models generally.
GLM 5.2's open-weights design offers practical advantages for cost-conscious and data-sensitive organizations. Zain Hasan, an AI engineer at Together AI who hosts the model on North American infrastructure, reported that GLM 5.2 excels at long-horizon tasks: earlier open-weights models often "lost the thread" after five to fifteen back-and-forth exchanges, but GLM 5.2 "could be using it for hours, and it would still have a coherent train of thought." David Nix, a principal software engineer at Denver-based MetaRouter, estimated that GLM 5.2 handles 10 to 20 percent of the work he sends to an LLM on a given day, particularly for front-end development where it comes "really close" to frontier models. Kacper Michalik, a software engineer at Kraków-based Screen Studio, achieved good results using GLM 5.2 to create website forms.
However, real-world experience has been uneven. Sai Kiran Myadaram, a software engineer at Bengaluru-based Indhic AI, signed up for Z.ai's subscription plan the week of launch and reported that his weekly token quota was "exhausted…in less than two to three days." He also encountered model hallucinations and over-planning on minor front-end fixes, which "messed up" his codebase, prompting him to revert to OpenAI's Codex. Michalik experienced occasional rate limits on the free plan but reported no significant quality issues. The cost advantage, while real, comes with a trade-off: GLM 5.2 appears most valuable as a supplementary tool for routine tasks rather than a replacement for frontier models on complex or critical work.
GLM 5.2's release reflects a widening competitive challenge for U.S. frontier AI labs. According to Stanford's AI Index, Chinese companies produced just over half as many notable AI models in 2025 as their U.S. counterparts, up from roughly a third in 2023. Z.ai's model is notable not only for its benchmark performance but for its pricing strategy: by undercutting U.S. rivals by an order of magnitude, it exposes a structural weakness in how many software companies manage AI spending. As Zain Hasan, an AI engineer at Together AI, observed, many firms lack token budgets and default to the most capable (and expensive) model when cost accountability is absent. This habit may have protected U.S. labs more than technical superiority alone.
However, GLM 5.2's practical limitations temper its threat. While it excels at long-horizon reasoning and front-end development, it significantly underperforms on harder agentic coding benchmarks—completing only 13 percent of SWE-Marathon tasks versus Opus 4.8's 26 percent. Real-world feedback also reveals rate-limiting issues and hallucinations on complex fixes, suggesting the model serves best as a supplementary tool rather than a complete frontier replacement. The open-weights design does provide a genuine advantage for organizations with data-sovereignty concerns: they can self-host and avoid sending information through Chinese infrastructure, a choice not available with most gated U.S. models.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion





Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack