AIToday
Large Language ModelsAI Business & IndustryITmedia AI+Published: Aug 27, 2026, 16:00 JST1 min read

Z.ai reveals 'Ox Alpha' is GLM-5.3-Flash, runs on China-made chips

Z.ai reveals 'Ox Alpha' is GLM-5.3-Flash, runs on China-made chips

Key takeaway

  • Z.ai revealed its anonymous model was GLM-5.3-Flash.

  • It ran on China-made chips, matching Nvidia GPU cost.

  • The model hit 17.5 trillion tokens on OpenRouter.

3 Key Points

  1. What happened

    Z.ai revealed on August 26 that the anonymous model "Ox Alpha" was actually its own GLM-5.3-Flash, and that inference during the validation period ran on a cluster of China-made AI chips. The model's weights were released under a commercial-use MIT license the same day.

  2. Why it matters

    GLM-5.3-Flash accumulated 17.5 trillion tokens on OpenRouter since August 20, ranking #1 in the weekly chart. Z.ai says the cluster of tens of thousands of China-made chips handled this demand at a per-token cost on par with Nvidia GPUs, arguing it shows China-made chips can support state-of-the-art inference at scale.

  3. What to watch

    The model has 320B total parameters and 18B active, supports text, image, and video input, and up to 1M token context. It scored 57 on the Artificial Analysis Intelligence Index, just 3 points behind the top model GLM-5.3, with about 7.5x lower cost per task.

Ask the AI about this article →

Context & Analysis

Z.ai has been offering GLM-5.3-Flash anonymously on OpenRouter since August 20, sparking speculation about its origin. The reveal on August 26 that it is a Z.ai model running on a cluster of tens of thousands of China-made chips is significant because it demonstrates that domestic hardware can serve frontier-level inference at a cost comparable to Nvidia GPUs. The model's performance—scoring 57 on the Artificial Analysis Intelligence Index, just 3 points behind the top model GLM-5.3—and its lower cost per task suggest that the combination of efficient architecture and China-made chips could offer a cost-effective alternative in the AI inference market. This development may encourage further investment in domestic chip ecosystems, though the article does not speculate on broader market impacts.

FAQ

What is the cost advantage of GLM-5.3-Flash?
The cost per task is about 7.5 times lower than the top model GLM-5.3, according to Artificial Analysis.
How does GLM-5.3-Flash achieve long-context efficiency?
It combines linear attention and sparse attention to compress past context and focus on key tokens, balancing accuracy and cost.

Get the latest Large Language Models news every morning

AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.

Free · takes 30 seconds · unsubscribe anytime

Ask AI

Ask AI anything about this article. Q&As are published on this page for other readers too.

Related Articles

Next articleRoboSense H1 revenue up 30.2%, robotics LiDAR sales surge 510.4%