Skip to content
context New models

The company xAI has released Grok 4.7 with a lower price but weaker benchmark results than Claude and GPT-6

only one source so far

The company xAI has released Grok 4.7 at $2/million input tokens and $6/million output tokens. On the Artificial Analysis index (v4.3.2), it achieves a score of 46, while Claude Fable 5.1 and GPT-6 lead with a score of 53 each. In the agentic coding test Terminal-Bench 4.0, it scores only 26 %.

The company xAI has launched Grok 4.7, which the company says is its most capable model yet for programming and knowledge work. According to xAI, it is built on a larger base model, trained with reinforcement learning for longer, and is intended to be better at verifying its own outputs. The price is $2 per million input tokens and $6 per million output tokens, which is closer to the prices of Chinese models than those of top-tier Western models.

On the independent Artificial Analysis Intelligence Index (version 4.3.2), which combines ten benchmarks, Grok 4.7 achieves a score of 46 and ranks in the middle of the pack. Claude Fable 5.1 and GPT-6 lead with a score of 53 each. The gap is even more pronounced on agentic coding tasks: in Terminal-Bench 4.0, Grok 4.7 achieves only 26 percent, while GPT-6 Astra achieves 60 percent and Claude Fable 5.1 achieves 55 percent. Even the cheaper model DeepSeek V4.1 Flash slightly outperforms it with 27 percent.

Grok 4.7 is available through Grok API, in the Cursor tool, and in Grok Build.

What changed

Why it matters

For developers and companies choosing a model for programming tasks, this is a concrete trade-off between price and performance: Grok 4.7 is cheaper than competing top-tier models, but according to documented benchmarks, it trails significantly in agentic coding, where even cheaper alternatives such as DeepSeek V4.1 Flash outperform it.

Release card

Grok 4.7

xAI

Price
$2 per million input tokens, $6 per million output tokens
Availability
Grok API, the Cursor tool, Grok Build
Documented measurements
  • Artificial Analysis Intelligence Index v4.3.2 46 bodů The aggregate score from ten benchmarks places the model in the middle of the pack, while Claude Fable 5.1 and GPT-6 achieve 53 points.
  • Terminal-Bench 4.0 26 % On the agentic coding task, it trails GPT-6 Astra (60 %) and Claude Fable 5.1 (55 %), as well as the cheaper model DeepSeek V4.1 Flash (27 %).
According to the sources, it is suitable for
  • programming (according to claims by xAI)
  • knowledge work (according to claims by xAI)
Documented limits
  • agentic coding tasks - it achieves only 26 % in Terminal-Bench 4.0, less than the cheaper model DeepSeek V4.1 Flash

In the Artificial Analysis Intelligence Index and Terminal-Bench 4.0 benchmarks, it trails Claude Fable 5.1 and GPT-6.

The card summarizes information from the article and any dated corrections, with a link to the original source. It is not our assessment of the model. You can find a dated profile with supporting materials at model page. Model selection and other announcements →

Two audiences, two different impacts

What this means

01

For individuals

A developer considering a model for agentic coding tasks should take into account that the lower price of Grok 4.7 goes hand in hand with a significantly lower score in Terminal-Bench 4.0 compared with Claude and GPT-6.

What to do Before deploying Grok 4.7 for programming tasks, compare its results in Terminal-Bench 4.0 with the performance of Claude and GPT-6 for the specific type of task.
More practical updates →
02

For a business

Companies planning to integrate Grok 4.7 into coding tools will get lower token costs, but according to documented benchmarks, they risk worse results on agentic coding tasks than with competing models.

Development
What to decide Test Grok 4.7 on your own agentic coding use case before deciding to switch because of the lower price, given the documented performance gap.
More business impacts →
benchmarky Claude GPT Grok kódování XAI

Check the original

Event sources

only one source so far · 1 publisher, 1 independent. We count feeds from the same owner only once.

1
The Decoder (daily AI news) independent context · first detected xAI launches Grok 4.7 at bargain prices, but benchmarks reveal a wide gap to Claude and GPT-6