Aakib Ansari.
Back to articles
News Brief

Z.AI Launches GLM-5.3-Flash: 320B Multimodal Open-Weight MoE Unveils Identity Behind Ox Alpha

Md Aakib Ansari
Md Aakib AnsariWeb Developer & AI Tools Reviewer
•3 min read•Model: GLM-5.3-Flash•Company: Z.AI
Z.AI Launches GLM-5.3-Flash: 320B Multimodal Open-Weight MoE Unveils Identity Behind Ox Alpha

On August 26, 2026, Beijing-based AI laboratory Z.AI (formerly Zhipu AI) officially released [GLM-5.3-Flash](/models/glm-5-3-flash), a 320-billion-parameter natively multimodal Mixture-of-Experts (MoE) model published under the permissive MIT license on Hugging Face.

The announcement formally confirms that GLM-5.3-Flash was the foundation behind the anonymous stealth model Ox Alpha, which topped usage charts and developer leaderboards on OpenRouter and OpenCode over the past week. By activating just 18 billion parameters per forward pass across a 1-million-token context window, the model is engineered to deliver frontier-tier reasoning and coding performance at a fraction of legacy serving costs.

On independent benchmark evaluations, Artificial Analysis ranked [GLM-5.3-Flash](/models/glm-5-3-flash) at a score of 57 on its Intelligence Index, placing it level with several closed commercial models while outperforming Z.AI's earlier GLM-5.2 base model. In community software engineering evaluations, the architecture previously demonstrated an 80% pass rate on DeepSWE sample suites during its stealth preview.

Z.AI has made full model weights immediately downloadable on Hugging Face (zai-org/GLM-5.3-Flash), with production API pricing set at $0.15 per million input tokens ($0.03 cached) and $0.50 per million output tokens on managed endpoints—discounted by 50% during a promotional launch window through early September.

The release marks a significant milestone in open-source multimodal intelligence, providing enterprise developers with an inspectable, MIT-licensed vision-language foundation that can be self-hosted across server-class clusters or routed via low-cost cloud APIs for high-volume agent workloads.

Frequently Asked Questions

Is GLM-5.3-Flash the same model as Ox Alpha?
Yes. Z.AI's official release confirms that GLM-5.3-Flash was the production checkpoint previewed anonymously under the codename 'Ox Alpha' on OpenRouter and OpenCode.
What are the parameter counts and license for GLM-5.3-Flash?
GLM-5.3-Flash features 320 billion total parameters with 18 billion active parameters per token, published as open weights under the MIT license with a 1M-token context window.

Related Articles