xAI · Strong performance at a low cost per task
Grok 4.5
A powerful general-purpose model also suited to programming, with a good balance of capabilities and cost. Suitable for tasks where the result can be verified.
- Costs
- Medium · approximately $0.31 per Intelligence Index task
- Speed
- Medium depending on the thoroughness setting
- Availability
- xAI's online service and products
- Input length
- Varies by specific version
- Inputs
- Text and, depending on the service used, image inputs
Data checked . Release date unverified.
Indicative capability profile
Where the model is strong
The five levels are our clear summary of the results below. They are not a ranking that applies to every task.
Ideal use
When to choose it
- Programming with tests and automated checks
- Solving complex problems on a limited budget
- A second model from a different provider for comparison
Usage boundaries
When to choose another model
- Outputs where facts cannot be verified
- Processes dependent on a single platform with no fallback option
Traceable supporting sources
Data and measurement sources
Distinguish between the manufacturer's documentation and the results of a specific test. Measurements also depend on the settings and the task set used.
Intelligence Index v4.1
54 pointsPerformance close to the best models at a low cost for the tested task.
Open original source ↗Coding Agent Index
76 points in Grok BuildThe result measures the model together with the tool it worked in.
Open original source ↗Trends over time
Related events from AI Radar
xAI made Grok 4.6 available in Amazon Bedrock
Developers building agents or tools for working with code on AWS can now try Grok 4.6 in Bedrock through Converse API with streaming and four reasoning effort levels, without needing an OpenAI-compatible client.
Companies running agentic or coding tools on AWS Bedrock gain another model option, with a choice of regional (USA) or global routing depending on data residency requirements and available capacity.
Next step
Compare the model using your actual task.
A benchmark narrows the selection. A short trial on your data determines the choice.