AIToday
Large Language ModelsOpen-Source AIAI Business & IndustryTHE DECODERPublished: Aug 27, 2026, 22:01 JST2 min read

GLM-5.3-Flash matches top models at a fraction of cost, runs without Nvidia

GLM-5.3-Flash matches top models at a fraction of cost, runs without Nvidia

Key takeaway

  • Z.ai's new GLM-5.3-Flash matches top models in intelligence but costs far less. It runs without Nvidia.

  • It served 100 trillion tokens daily on Chinese chips.

  • This challenges the CUDA moat.

3 Key Points

  1. What happened

    Z.ai released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. It scores 57 points on the Intelligence Index, level with GPT-5.6 Terra and Muse Spark 1.2, and just three points behind the larger GLM-5.3.

  2. Why it matters

    The model costs 0.09 dollars per task on the index, versus 0.68 dollars for GLM-5.3 — about 7.5 times cheaper. It also runs on Chinese AI chips, with SemiAnalysis reporting it served 100 trillion tokens a day, testing the CUDA moat.

  3. What to watch

    On the API, it costs 0.15 dollars per million input tokens and 0.50 dollars per million output tokens, about ten percent of GLM-5.3's price. On agentic tasks, it matches GLM-5.3 and Grok 4.6 on GDPval-AA v2, trailing only Claude Opus 5.

Ask the AI about this article →

Context & Analysis

Z.ai's new GLM-5.3-Flash stands out not just for its performance but for its cost structure. The model reaches 57 points on the Intelligence Index, only three behind the larger GLM-5.3, while costing 0.09 dollars per task — roughly 7.5 times cheaper. This places it on the Pareto frontier of intelligence and cost, a position that adds to the recent pressure Chinese models have put on Western providers.

The infrastructure angle is equally significant. Z.ai ran the model's pre-launch traffic on Chinese AI chips, serving 100 trillion tokens a day as reported by SemiAnalysis. The company says its hardware efficiency matches common Nvidia GPUs. Because CUDA, Nvidia's programming layer, has been nearly 20 years of tuned integration, switching chips usually requires redoing that work. Z.ai instead built its own serving software on SGLang, breaking processing into independently scaling stages, which tripled throughput on the same hardware.

FAQ

How much does GLM-5.3-Flash cost on the API?
On Z.ai's API, it costs 0.15 dollars per million input tokens and 0.50 dollars per million output tokens, about a tenth of GLM-5.3's price.
How does GLM-5.3-Flash perform on agentic tasks?
On GDPval-AA v2, it hits an Elo score of about 1770, matching GLM-5.3 and Grok 4.6, and trails only Claude Opus 5.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleNvidia to acquire Hugging Face for $12.9 billion: report