Skip to content
important New models verified update

Z.ai runs GLM-5.3-Flash on 100 000 Chinese AI chips instead of Nvidia GPUs

confirmed by 2 independent sources updated August 27, 2026

Z.ai confirmed that GLM-5.3-Flash (320 billion parameters) runs on 100 000 Chinese AI chips instead of Nvidia GPUs. Released on 26 August 2026 under the MIT license, the model achieves performance close to GLM-5.3 at roughly one-seventh of the price.

Z.ai confirmed that 100 000 Chinese-made chips, rather than Nvidia graphics cards, handle GLM-5.3-Flash operation for online queries. According to the company, its own software stack built on SGLang achieved efficiency on these chips comparable to standard Nvidia GPUs and tripled throughput compared with the first attempt on the same hardware; an agent built on GLM-5.3 also contributed to the optimization. SemiAnalysis reported that the model can serve 100 trillion tokens per day this way. The move follows efforts by China to reduce dependence on Nvidia - in September 2025, the Chinese government ordered Alibaba and ByteDance to stop buying Nvidia AI chips, and in November it banned foreign AI chips in state-funded data centers.

Z.ai initially tested GLM-5.3-Flash under the codename Ox Alpha from 20 August 2026, officially releasing it on 26 August 2026 as an open-weight model under the MIT license. It is the first natively multimodal model in the GLM-5 series, with 320 billion parameters, of which 18 billion are active, and a context window of 1 million tokens; the weights are available on Hugging Face. According to Z.ai, the model focuses on long contexts, visually guided agentic tasks and code synthesis.

According to measurements by Artificial Analysis, the model achieves 57 points in the Intelligence Index at the maximum reasoning setting, just three points below the larger GLM-5.3 (60 points) and comparable to GPT-5.6 Terra and Muse Spark 1.2. The cost per task according to the same index is 0.09 dollars compared with 0.68 dollars for GLM-5.3, roughly 7.5 times lower. Through the API from Z.ai, the model costs 0.075 dollars per million input tokens and 0.25 dollars per million output tokens until 9 September 2026, after which prices rise to 0.15 and 0.50 dollars per million tokens, respectively. For comparison, GPT-5.6 Luna from OpenAI costs 0.20 dollars per million input tokens and 2 dollars per million output tokens, while Claude Opus 5 from Anthropic costs 5 dollars per million input tokens and per million output tokens. On the GDPval-AA v2 benchmark, GLM-5.3-Flash achieves an Elo score of around 1770, equal to GLM-5.3 and Grok 4.6, and trails only Claude Opus 5. According to Artificial Analysis, approximately 90 percent of the output tokens from the model go toward reasoning, making it less token-efficient.

What changed

Why it matters

Deploying the model exclusively on Chinese chips shows that, according to the cited analysts, Chinese hardware may already be sufficient for companies for inference - unlike training the most powerful models - reducing dependence on Nvidia among Chinese providers. For users and companies outside China, this means further pressure on prices from Western providers and the ability to route tasks between multiple models based on their price/performance ratio, while questions about data security and regulation for services hosted in China remain open.

Release card

GLM-5.3-Flash

Z.ai

open weights
Specifications
320 billion parameters in total, 18 billion active
Context
1 million tokens
Inputs
natively multimodal model; supports visually guided agentic tasks and code synthesis
Licence
MIT
Price
0.075 dollars per million input tokens and 0.25 dollars per million output tokens until 9 September 2026, then 0.15 dollars per million input tokens and 0.50 dollars per million output tokens (API from Z.ai)
Availability
The model weights are available on Hugging Face under the MIT license, and the model is also available through the API from Z.ai.
Documented measurements
According to the sources, it is suitable for
  • long contexts
  • visually guided agentic tasks
  • code synthesis
Documented limits
  • lower token efficiency, approximately 90 % of output tokens go toward reasoning

The model achieves performance close to the larger GLM-5.3 model (57 points versus 60 in the Intelligence Index) at a significantly lower price; on the GDPval-AA v2 benchmark, it is comparable to GLM-5.3 and Grok 4.6, but trails Claude Opus 5.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. It does not yet have a dedicated editorial profile. Model selection and other announcements →

What was added since the original report

Verified updates

  1. New verified information

    The model was predicted on 20 August under the codename 'Ox Alpha', officially released on 26 August 2026; It is being deployed on 100 000 Chinese AI chips instead of Nvidia; Temporary price of 0.075 USD per million input tokens (until 9 September), then 0.15 USD; output tokens 0.25 USD/million; Focused on long contexts, agentic tasks and code synthesis

    • The model was predicted on 20 August under the codename 'Ox Alpha', officially released on 26 August 2026
    • It is being deployed on 100 000 Chinese AI chips instead of Nvidia
    • Temporary price of 0.075 USD per million input tokens (until 9 September), then 0.15 USD; output tokens 0.25 USD/million
    • Focused on long contexts, agentic tasks and code synthesis

Two audiences, two different impacts

What this means

01

For individuals

Developers have access to an open-weight model with performance close to the more expensive GLM-5.3, but at a significantly lower price than GPT-5.6 Luna or Claude Opus 5, also available for download under the MIT license.

What to do Consider trying GLM-5.3-Flash through an API or OpenRouter for non-critical agentic or coding tasks and comparing its price/performance ratio with the model currently in use.
More practical updates →
02

For a business

The availability of a powerful model running on Chinese chips at a fraction of the price allows companies to design AI systems that route tasks between multiple providers based on price and performance instead of relying on a single vendor; however, according to the cited analysts, questions about data security, regulation and long-term reliability remain open for services hosted in China…

Strategy
What to decide Before any potential deployment of GLM-5.3-Flash or similar Chinese models, review requirements for data security, regulation and the ability to route tasks between multiple providers (multi-model routing).
More business impacts →
Chinese AI chips Chinese chips cost efficiency GLM-5.3-Flash inference competition with Nvidia open-source open-source models Z.ai

Check the original

Event sources

confirmed by 2 independent sources · 2 publishers, 2 independent. We count feeds from the same owner only once.

2
The Decoder (daily AI news) independent context · first detected GLM-5.3-Flash matches top models at a fraction of the cost, and runs without Nvidia AI Business independent context Z.AI's Use of Chinese Chips for New Model is About Optimization