Z.ai runs GLM-5.3-Flash on 100 000 Chinese AI chips instead of Nvidia GPUs
Z.ai confirmed that GLM-5.3-Flash (320 billion parameters) runs on 100 000 Chinese AI chips instead of Nvidia GPUs. Released on 26 August 2026 under the MIT license, the model achieves performance close to GLM-5.3 at roughly one-seventh of the price.
Z.ai confirmed that 100 000 Chinese-made chips, rather than Nvidia graphics cards, handle GLM-5.3-Flash operation for online queries. According to the company, its own software stack built on SGLang achieved efficiency on these chips comparable to standard Nvidia GPUs and tripled throughput compared with the first attempt on the same hardware; an agent built on GLM-5.3 also contributed to the optimization. SemiAnalysis reported that the model can serve 100 trillion tokens per day this way. The move follows efforts by China to reduce dependence on Nvidia - in September 2025, the Chinese government ordered Alibaba and ByteDance to stop buying Nvidia AI chips, and in November it banned foreign AI chips in state-funded data centers.
Z.ai initially tested GLM-5.3-Flash under the codename Ox Alpha from 20 August 2026, officially releasing it on 26 August 2026 as an open-weight model under the MIT license. It is the first natively multimodal model in the GLM-5 series, with 320 billion parameters, of which 18 billion are active, and a context window of 1 million tokens; the weights are available on Hugging Face. According to Z.ai, the model focuses on long contexts, visually guided agentic tasks and code synthesis.
According to measurements by Artificial Analysis, the model achieves 57 points in the Intelligence Index at the maximum reasoning setting, just three points below the larger GLM-5.3 (60 points) and comparable to GPT-5.6 Terra and Muse Spark 1.2. The cost per task according to the same index is 0.09 dollars compared with 0.68 dollars for GLM-5.3, roughly 7.5 times lower. Through the API from Z.ai, the model costs 0.075 dollars per million input tokens and 0.25 dollars per million output tokens until 9 September 2026, after which prices rise to 0.15 and 0.50 dollars per million tokens, respectively. For comparison, GPT-5.6 Luna from OpenAI costs 0.20 dollars per million input tokens and 2 dollars per million output tokens, while Claude Opus 5 from Anthropic costs 5 dollars per million input tokens and per million output tokens. On the GDPval-AA v2 benchmark, GLM-5.3-Flash achieves an Elo score of around 1770, equal to GLM-5.3 and Grok 4.6, and trails only Claude Opus 5. According to Artificial Analysis, approximately 90 percent of the output tokens from the model go toward reasoning, making it less token-efficient.
Why it matters
Deploying the model exclusively on Chinese chips shows that, according to the cited analysts, Chinese hardware may already be sufficient for companies for inference - unlike training the most powerful models - reducing dependence on Nvidia among Chinese providers. For users and companies outside China, this means further pressure on prices from Western providers and the ability to route tasks between multiple models based on their price/performance ratio, while questions about data security and regulation for services hosted in China remain open.
Release card
GLM-5.3-Flash
Z.ai
- Specifications
- 320 billion parameters in total, 18 billion active
- Context
- 1 million tokens
- Inputs
- natively multimodal model; supports visually guided agentic tasks and code synthesis
- Licence
- MIT
- Price
- 0.075 dollars per million input tokens and 0.25 dollars per million output tokens until 9 September 2026, then 0.15 dollars per million input tokens and 0.50 dollars per million output tokens (API from Z.ai)
- Availability
- The model weights are available on Hugging Face under the MIT license, and the model is also available through the API from Z.ai.
- Artificial Analysis Intelligence Index 57 bodů při maximálním nastavení uvažování The score is three points lower than that of GLM-5.3 (60 points) and comparable to GPT-5.6 Terra and Muse Spark 1.2.
- Artificial Analysis - cena za úlohu (Intelligence Index) 0,09 dolaru za úlohu Approximately 7.5 times lower cost per task than GLM-5.3 (0.68 dollars).
- GDPval-AA v2 Elo skóre přibližně 1770 Equal to GLM-5.3 and Grok 4.6; the model trails only Claude Opus 5.
- long contexts
- visually guided agentic tasks
- code synthesis
- lower token efficiency, approximately 90 % of output tokens go toward reasoning
The model achieves performance close to the larger GLM-5.3 model (57 points versus 60 in the Intelligence Index) at a significantly lower price; on the GDPval-AA v2 benchmark, it is comparable to GLM-5.3 and Grok 4.6, but trails Claude Opus 5.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →
What was added since the original report
Verified updates
-
The model was predicted on 20 August under the codename 'Ox Alpha', officially released on 26 August 2026; It is being deployed on 100 000 Chinese AI chips instead of Nvidia; Temporary price of 0.075 USD per million input tokens (until 9 September), then 0.15 USD; output tokens 0.25 USD/million; Focused on long contexts, agentic tasks and code synthesis
- The model was predicted on 20 August under the codename 'Ox Alpha', officially released on 26 August 2026
- It is being deployed on 100 000 Chinese AI chips instead of Nvidia
- Temporary price of 0.075 USD per million input tokens (until 9 September), then 0.15 USD; output tokens 0.25 USD/million
- Focused on long contexts, agentic tasks and code synthesis
Two audiences, two different impacts
What this means
For individuals
Developers have access to an open-weight model with performance close to the more expensive GLM-5.3, but at a significantly lower price than GPT-5.6 Luna or Claude Opus 5, also available for download under the MIT license.
For a business
The availability of a powerful model running on Chinese chips at a fraction of the price allows companies to design AI systems that route tasks between multiple providers based on price and performance instead of relying on a single vendor; however, according to the cited analysts, questions about data security, regulation and long-term reliability remain open for services hosted in China…
StrategyCheck the original
Event sources
confirmed by 2 independent sources · 2 publishers, 2 independent. We count feeds from the same owner only once.