Z.ai · A model to run on your own server
GLM-5.2
Model files can be downloaded and run on your own server, giving organizations greater control over their data. However, this requires very powerful hardware.
- Costs
- Variable · online service $1.40 per million input tokens and $4.40 per million output tokens; self-hosting requires powerful hardware
- Speed
- Medium to slow; generates large amounts of text
- Availability
- It can be downloaded, modified and also used through several online services
- Input length
- Very long input · up to 1 million tokens (small pieces of text)
- Inputs
- Primarily text; depends on how it is run
Data checked . Release date unverified.
Indicative capability profile
Where the model is strong
The five levels are our clear summary of the results below. They are not a ranking that applies to every task.
Ideal use
When to choose it
- Running on your own or a rented server
- Working with sensitive data that must remain under the company's control
- Multi-step tasks with the option to modify the model
Usage boundaries
When to choose another model
- Running on a standard personal computer
- Teams that do not want to manage powerful servers
Traceable supporting sources
Data and measurement sources
Distinguish between the manufacturer's documentation and the results of a specific test. Measurements also depend on the settings and the task set used.
Intelligence Index v4.1
51 pointsThe highest-rated downloadable model in the measurement from 16 June 2026.
Open original source ↗Terminal-Bench · SciCode
78 % · 50 %Very good result in coding and autonomous terminal work.
Open original source ↗Trends over time
Related events from AI Radar
Kimi K3 from Moonshot AI is newly available on the Amazon Bedrock platform
New: Kimi K3 is now available on Amazon Bedrock as a new distribution channel; The model supports prompt caching on Bedrock; Data in AWS Bedrock is protected by the AWS data boundary without being shared with the provider; The Bedrock deployment optimizes the model for long coding and knowledge workflows
Developers can try Kimi K3 through the API or through tools such as OpenCode for agentic programming tasks with long context, but must account for a higher price per token than with previous Kimi models.
Companies can deploy Kimi K3 through Amazon Bedrock with data protection from AWS without needing their own GPU infrastructure, or host the model themselves on instances with 8 NVIDIA B300 GPUs – but both options involve higher costs than earlier open Kimi models.
Frontier Red Team at Anthropic: GLM-5.3 and Claude Mythos Preview have crossed a threshold in binary exploitation capabilities
Companies working on security, red-teaming or compliance for AI systems are receiving a signal that the latest models have crossed a threshold in offensive cyber capabilities (control flow hijack), which is relevant to assessing the risks of AI misuse and setting security policies around the deployment of these models.
Next step
Compare the model using your actual task.
A benchmark narrows the selection. A short trial on your data determines the choice.