
Z.ai revealed its anonymous model was GLM-5.3-Flash.
It ran on China-made chips, matching Nvidia GPU cost.
The model hit 17.5 trillion tokens on OpenRouter.
What happened
Z.ai revealed on August 26 that the anonymous model "Ox Alpha" was actually its own GLM-5.3-Flash, and that inference during the validation period ran on a cluster of China-made AI chips. The model's weights were released under a commercial-use MIT license the same day.
Why it matters
GLM-5.3-Flash accumulated 17.5 trillion tokens on OpenRouter since August 20, ranking #1 in the weekly chart. Z.ai says the cluster of tens of thousands of China-made chips handled this demand at a per-token cost on par with Nvidia GPUs, arguing it shows China-made chips can support state-of-the-art inference at scale.
What to watch
The model has 320B total parameters and 18B active, supports text, image, and video input, and up to 1M token context. It scored 57 on the Artificial Analysis Intelligence Index, just 3 points behind the top model GLM-5.3, with about 7.5x lower cost per task.
Ask the AI about this article →
Z.ai has been offering GLM-5.3-Flash anonymously on OpenRouter since August 20, sparking speculation about its origin. The reveal on August 26 that it is a Z.ai model running on a cluster of tens of thousands of China-made chips is significant because it demonstrates that domestic hardware can serve frontier-level inference at a cost comparable to Nvidia GPUs. The model's performance—scoring 57 on the Artificial Analysis Intelligence Index, just 3 points behind the top model GLM-5.3—and its lower cost per task suggest that the combination of efficient architecture and China-made chips could offer a cost-effective alternative in the AI inference market. This development may encourage further investment in domestic chip ecosystems, though the article does not speculate on broader market impacts.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Plaud Inc. introduced the Plaud One Explorer Edition, a pair of earbuds and a charging case that serve as a mo…
Chinese AI startup MiniMax Group reported a 283.1% increase in revenue, from US$30.4 million for the six month…

Nvidia's GPU roadmap is pushing 800VDC as a next-generation breakthrough, driving battery backup modules (BBUs…

Sophus IT Solutions, an AI native engineering and professional services firm, has been named an OpenAI Select…

Z.ai released GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series

Google has launched Gemini 3.5 Transcribe, a speech-to-text model for real-time transcription that auto-correc…
