
GLM-5.3, a new open-source language model from Chinese startup z.ai, has launched on API at $1.40 per million input tokens and $4.40 per million output tokens—the same price as its predecessor.
The model drew attention for detecting a previously unknown vulnerability in Cursor, and z.ai plans eventually to release the weights openly, though timing and licensing details remain unclear.
What happened
GLM-5.3, a new open-source language model from Chinese startup z.ai, is now available via API following its debut last week. Developers with a GLM Coding Plan subscription can access it through an OpenAI Chat Completions-compatible protocol. Pricing matches the previous GLM-5.2 generation: $1.40 per million input tokens and $4.40 per million output tokens, with cached input at $0.26 per million tokens and free cached-input storage for a limited time.
Why it matters
The model gained attention for advanced cyber capabilities that reportedly uncovered a previously undetected vulnerability in Cursor. By keeping API pricing flat between generations, z.ai removes a cost barrier for developers considering migration from competitors, potentially lowering adoption friction during the transition.
What to watch
Z.ai said it plans to make the model's weights openly available, though a precise date and licensing terms have not yet been announced. This open-source release could further broaden access beyond API consumers.
Ask the AI about this article →
GLM-5.3 arrives at a moment when z.ai is working to establish its open-source language model in a competitive landscape dominated by established players. The model's reported ability to detect a previously undetected vulnerability in Cursor—a code editor—signals meaningful capability in specialized domains, which can drive adoption among developers building security-focused or code-related applications.
By maintaining the same per-token pricing as GLM-5.2, z.ai removes a potential friction point for developers evaluating whether to migrate from other models. In a market where token costs influence the unit economics of AI applications, pricing parity between generations lowers the switching barrier. The addition of free cached-input storage (for a limited time) further sweetens the offer for high-volume or repeated-query workloads.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Anthropic is privately hoping to file for its initial public offering by the end of this month, targeting a ra…
Broadcom is reportedly seeking to borrow up to $100 billion in debt financing to support growth efforts at Ant…
As AI technology matures, the bottleneck in the industry is moving beyond semiconductor constraints like GPUs…

Elice Group, a South Korean AI infrastructure provider, announced the launch of the country's first AI data ce…

On August 12, AT&T's Chief Data and AI Officer said OpenAI models power about 25% of the telecom's total AI us…

On August 11, IBM announced a multi-year $240 million agreement with Together AI to deploy NVIDIA HGX B300 sys…
