xAI · Code and working with information
Grok 4.7
A model for coding and multistep tasks involving text and images.
- Costs
- $2 input / $6 output per million tokens below 200 thousand input tokens; above this threshold, $4 / $12.
- Speed
- We have not independently measured speed on a comparable basis.
- Availability
- xAI API, Cursor and Grok Build
- Input length
- 500 thousand tokens.
- Inputs
- Text and images → text
Data checked . Published September 21, 2026. Specifications and prices are provided by the vendor; usage recommendations come from the editorial team. We do not yet have our own comparative test of this version.
Ideal use
When to choose it
- Code development and fixes
- Analysis of text and image materials
Usage boundaries
When to choose another model
- The Fast variant is not available in the public xAI API.
Traceable supporting sources
Data and measurement sources
Distinguish between the manufacturer's documentation and the results of a specific test. Measurements also depend on the settings and the task set used.
Vendor documentation
Specifications, availability, and termsManufacturer data verified as of the review date. Recommended use is an editorial interpretation.
Open original source ↗Efficiency in practice
With what Grok 4.7 combine
Quick code edits with a second check
One system makes the change, another looks for blind spots, and tests decide.
A second model will not help if both receive the same incorrect assumption without evidence.Trends over time
Related events from AI Radar
Grok 4.7 from xAI available on Amazon Bedrock
Developers working with Amazon Bedrock have another model available for coding and long-running agentic tasks. According to the data, it is worth deliberately choosing the reasoning effort level because of the noticeable difference in token consumption and therefore in response speed and cost.
Companies building agentic systems or automation on Amazon Bedrock gain another frontier model with a long context window and self-verification that, according to Artificial Analysis, reduces error rates on long tasks but consumes roughly twice as many output tokens as the previous version at the highest reasoning effort level, increasing operating costs.
GitHub Copilot makes Grok 4.7 from xAI available
Developers with a paid Copilot SKU (Pro, Pro+, Max, Business, Enterprise) gain a new model option focused on agentic coding in the model picker once the gradual rollout reaches them.
Companies with Copilot Business or Enterprise must manage access to the new model through the model policy and account for the fact that new models are automatically enabled by default, which affects costs when billing is based on the provider's list prices.
The company xAI has released Grok 4.7 with a lower price but weaker benchmark results than Claude and GPT-6
A developer considering a model for agentic coding tasks should take into account that the lower price of Grok 4.7 goes hand in hand with a significantly lower score in Terminal-Bench 4.0 compared with Claude and GPT-6.
Companies planning to integrate Grok 4.7 into coding tools will get lower token costs, but according to documented benchmarks, they risk worse results on agentic coding tasks than with competing models.
OpenAI halved prices for the GPT-6 Sol and Luna models; independent analysis does not confirm a performance improvement
New: Anthropic released Claude Opus 5.5 on the same day (2026-09-23), roughly an hour after OpenAI; Claude Opus 5.5 is 20% cheaper than the previous Opus 5.0 ($4/$20 vs $5/$25); Opus 5.5 improved communication style and token efficiency; The model releases launch a 'price war' between the two largest AI companies; GPT-6 Luna and Sol are part of a broader strategy of price competition, not an isolated release
Developers working with the API from OpenAI can reduce inference costs by up to half while maintaining the same usage pattern (Sol for more complex tasks, Luna for high-volume simple tasks), but independent tests suggest that a lower price may not mean higher or even the same output quality in all cases.
Businesses gain a cheaper option for reasoning models for high-volume or repetitive tasks (summaries, data extraction, code), but according to an independent analysis, actual output quality has not improved compared with the previous generation and has instead declined on some knowledge work tasks, so deployment decisions should not rely solely on the manufacturer's claims about price/performance.
Next step
Compare the model using your actual task.
A benchmark narrows the selection. A short trial on your data determines the choice.