AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryVentureBeat AIPublished: Aug 19, 2026, 13:01 JST2 min read

GLM-5.3 API launches at $1.4/$4.4 per million tokens

GLM-5.3 API launches at $1.4/$4.4 per million tokens

Key takeaway

  • GLM-5.3, a new open-source language model from Chinese startup z.ai, has launched on API at $1.40 per million input tokens and $4.40 per million output tokens—the same price as its predecessor.

  • The model drew attention for detecting a previously unknown vulnerability in Cursor, and z.ai plans eventually to release the weights openly, though timing and licensing details remain unclear.

3 Key Points

  1. What happened

    GLM-5.3, a new open-source language model from Chinese startup z.ai, is now available via API following its debut last week. Developers with a GLM Coding Plan subscription can access it through an OpenAI Chat Completions-compatible protocol. Pricing matches the previous GLM-5.2 generation: $1.40 per million input tokens and $4.40 per million output tokens, with cached input at $0.26 per million tokens and free cached-input storage for a limited time.

  2. Why it matters

    The model gained attention for advanced cyber capabilities that reportedly uncovered a previously undetected vulnerability in Cursor. By keeping API pricing flat between generations, z.ai removes a cost barrier for developers considering migration from competitors, potentially lowering adoption friction during the transition.

  3. What to watch

    Z.ai said it plans to make the model's weights openly available, though a precise date and licensing terms have not yet been announced. This open-source release could further broaden access beyond API consumers.

Ask the AI about this article →

Context & Analysis

GLM-5.3 arrives at a moment when z.ai is working to establish its open-source language model in a competitive landscape dominated by established players. The model's reported ability to detect a previously undetected vulnerability in Cursor—a code editor—signals meaningful capability in specialized domains, which can drive adoption among developers building security-focused or code-related applications.

By maintaining the same per-token pricing as GLM-5.2, z.ai removes a potential friction point for developers evaluating whether to migrate from other models. In a market where token costs influence the unit economics of AI applications, pricing parity between generations lowers the switching barrier. The addition of free cached-input storage (for a limited time) further sweetens the offer for high-volume or repeated-query workloads.

FAQ

What is the pricing for GLM-5.3 on the API?
Input tokens cost $1.40 per million, output tokens cost $4.40 per million, cached input costs $0.26 per million tokens, and cached-input storage is currently free for a limited time.
Who can currently access GLM-5.3 via the API?
Developers who previously subscribed to a GLM Coding Plan are currently limited to the OpenAI Chat Completions-compatible protocol.
Will the model weights be made public?
Z.ai said it plans to make the model's weights openly available, but a precise date and licensing remain to be seen.
VentureBeat AIRead Original Article

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleGoogle uses AI to guide flights around contrails, cutting aviation climate impact

The AI news that matters, in one minute each morning.

Sign up free