The company xAI has released Grok 4.7 with a lower price but weaker benchmark results than Claude and GPT-6
The company xAI has released Grok 4.7 at $2/million input tokens and $6/million output tokens. On the Artificial Analysis index (v4.3.2), it achieves a score of 46, while Claude Fable 5.1 and GPT-6 lead with a score of 53 each. In the agentic coding test Terminal-Bench 4.0, it scores only 26 %.
The company xAI has launched Grok 4.7, which the company says is its most capable model yet for programming and knowledge work. According to xAI, it is built on a larger base model, trained with reinforcement learning for longer, and is intended to be better at verifying its own outputs. The price is $2 per million input tokens and $6 per million output tokens, which is closer to the prices of Chinese models than those of top-tier Western models.
On the independent Artificial Analysis Intelligence Index (version 4.3.2), which combines ten benchmarks, Grok 4.7 achieves a score of 46 and ranks in the middle of the pack. Claude Fable 5.1 and GPT-6 lead with a score of 53 each. The gap is even more pronounced on agentic coding tasks: in Terminal-Bench 4.0, Grok 4.7 achieves only 26 percent, while GPT-6 Astra achieves 60 percent and Claude Fable 5.1 achieves 55 percent. Even the cheaper model DeepSeek V4.1 Flash slightly outperforms it with 27 percent.
Grok 4.7 is available through Grok API, in the Cursor tool, and in Grok Build.
Why it matters
For developers and companies choosing a model for programming tasks, this is a concrete trade-off between price and performance: Grok 4.7 is cheaper than competing top-tier models, but according to documented benchmarks, it trails significantly in agentic coding, where even cheaper alternatives such as DeepSeek V4.1 Flash outperform it.
Release card
Grok 4.7
xAI
- Price
- $2 per million input tokens, $6 per million output tokens
- Availability
- Grok API, the Cursor tool, Grok Build
- Artificial Analysis Intelligence Index v4.3.2 46 bodů The aggregate score from ten benchmarks places the model in the middle of the pack, while Claude Fable 5.1 and GPT-6 achieve 53 points.
- Terminal-Bench 4.0 26 % On the agentic coding task, it trails GPT-6 Astra (60 %) and Claude Fable 5.1 (55 %), as well as the cheaper model DeepSeek V4.1 Flash (27 %).
- programming (according to claims by xAI)
- knowledge work (according to claims by xAI)
- agentic coding tasks - it achieves only 26 % in Terminal-Bench 4.0, less than the cheaper model DeepSeek V4.1 Flash
In the Artificial Analysis Intelligence Index and Terminal-Bench 4.0 benchmarks, it trails Claude Fable 5.1 and GPT-6.
The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. You can find a dated profile with supporting materials at model page. Model selection and other announcements →
Two audiences, two different impacts
What this means
For individuals
A developer considering a model for agentic coding tasks should take into account that the lower price of Grok 4.7 goes hand in hand with a significantly lower score in Terminal-Bench 4.0 compared with Claude and GPT-6.
For a business
Companies planning to integrate Grok 4.7 into coding tools will get lower token costs, but according to documented benchmarks, they risk worse results on agentic coding tasks than with competing models.
DevelopmentCheck the original
Event sources
only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.