
Z.ai has open-sourced GLM-5.3-Flash, an LLM that costs ten times less to run than its predecessor.
It scored highest on the GDPval-AA v2 benchmark.
The model's weights are available on Hugging Face.
What happened
Z.ai Co. released the code for GLM-5.3-Flash, an LLM with 320 billion parameters that activates 18 billion per prompt. It first appeared last week on OpenRouter Inc. as an unnamed free model called Ox Alpha, which drew industry attention.
Why it matters
The model costs ten times less to run than Z.ai's previous-generation LLM. It uses sparse attention and linear attention to cut processing power and memory, and it scored highest on the GDPval-AA v2 benchmark against Claude Opus 4.8, GPT-5.6 Terra and Gemini 3.7 Flash.
What to watch
GLM-5.3-Flash accepts up to 1 million tokens of input (text, images, video) and generates up to 131,072 tokens. Its weights are available on Hugging Face.
Ask the AI about this article →
Z.ai's open-sourcing of GLM-5.3-Flash follows the model's debut on OpenRouter as 'Ox Alpha' without a named developer, a move that sparked speculation and attention. The company has now confirmed its authorship and released the weights publicly on Hugging Face.
The release emphasizes cost efficiency, a key factor for businesses deploying AI at scale. The company states the model costs ten times less to run than its prior generation, achieved through architectural changes including sparse and linear attention mechanisms that reduce processing and memory demands.
Benchmark results place GLM-5.3-Flash competitively, with the highest score on GDPval-AA v2 against Claude Opus 4.8, GPT-5.6 Terra, and Gemini 3.7 Flash, and second place on AutomationBench. Training used a 30 trillion token dataset and a technology called mHC to reduce gradient distortion.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
Ask AI anything about this article. Q&As are published on this page for other readers too.
Deep Cogito Inc. raised $43 million in a Series A round led by TQ Ventures, with participation from Benchmark…
Beijing introduced China's first dedicated "AI4Chip" policy, extending AI into semiconductor production across…

Many cloud-based AI services, including ChatGPT, use input data for AI training by default, even on paid perso…

Daily Dose of Data Science built an AI workflow using Mistral OCR 4 that reads every chart in a scientific pap…

Alibaba's Qwen team released Qwen3.8-Flash-Next, an open-weights multimodal MoE model that also serves as an e…

Meta's Project OT, an internal plan to use AI agents to replace workers, would have reduced headcount by about…
