
Z.ai's new GLM-5.3-Flash matches top models in intelligence but costs far less. It runs without Nvidia.
It served 100 trillion tokens daily on Chinese chips.
This challenges the CUDA moat.
What happened
Z.ai released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. It scores 57 points on the Intelligence Index, level with GPT-5.6 Terra and Muse Spark 1.2, and just three points behind the larger GLM-5.3.
Why it matters
The model costs 0.09 dollars per task on the index, versus 0.68 dollars for GLM-5.3 — about 7.5 times cheaper. It also runs on Chinese AI chips, with SemiAnalysis reporting it served 100 trillion tokens a day, testing the CUDA moat.
What to watch
On the API, it costs 0.15 dollars per million input tokens and 0.50 dollars per million output tokens, about ten percent of GLM-5.3's price. On agentic tasks, it matches GLM-5.3 and Grok 4.6 on GDPval-AA v2, trailing only Claude Opus 5.
Ask the AI about this article →
Z.ai's new GLM-5.3-Flash stands out not just for its performance but for its cost structure. The model reaches 57 points on the Intelligence Index, only three behind the larger GLM-5.3, while costing 0.09 dollars per task — roughly 7.5 times cheaper. This places it on the Pareto frontier of intelligence and cost, a position that adds to the recent pressure Chinese models have put on Western providers.
The infrastructure angle is equally significant. Z.ai ran the model's pre-launch traffic on Chinese AI chips, serving 100 trillion tokens a day as reported by SemiAnalysis. The company says its hardware efficiency matches common Nvidia GPUs. Because CUDA, Nvidia's programming layer, has been nearly 20 years of tuned integration, switching chips usually requires redoing that work. Z.ai instead built its own serving software on SGLang, breaking processing into independently scaling stages, which tripled throughput on the same hardware.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Plaud Inc. introduced the Plaud One Explorer Edition, a pair of earbuds and a charging case that serve as a mo…
Chinese AI startup MiniMax Group reported a 283.1% increase in revenue, from US$30.4 million for the six month…

Nvidia's GPU roadmap is pushing 800VDC as a next-generation breakthrough, driving battery backup modules (BBUs…

Sophus IT Solutions, an AI native engineering and professional services firm, has been named an OpenAI Select…

Anthropic has agreed to rent compute capacity from British cloud startup Nscale for around $45 billion over si…

Google has launched Gemini 3.5 Transcribe, a speech-to-text model for real-time transcription that auto-correc…
