DeepSeek has released V4 Flash 0731 with performance close to GPT-5.6 Luna at a significantly lower price
DeepSeek has released V4 Flash 0731 with a score of 50 points in the Artificial Analysis Intelligence Index, one point below GPT-5.6 Luna, but at a price about 60 % lower. The weights are under the MIT license on Hugging Face, with a price of 0.14 $/0.27 $ per million tokens.
DeepSeek has released a new V4 Flash model designated „0731", an update to its affordable model. According to the Artificial Analysis Intelligence Index, the model scores 50 points, ten more than the previous V4 Flash version from April 2026. The score is just one point lower than that of GPT-5.6 Luna from OpenAI, yet according to the sources, a task with V4 Flash is approximately 60 percent cheaper than with GPT-5.6 Luna, even after its 80 percent price reduction. According to Artificial Analysis, the model also outperforms MiniMax M3 (a model with 428 billion parameters) in the Intelligence Index.
The price is 0.14 dollars per million input tokens and 0.27 dollars per million output tokens. DeepSeek reports a cache discount of 98 percent, above the industry standard of 90 percent. According to the company, the model uses 12 percent fewer tokens than the previous version. The architecture remains the same: 284 billion parameters in total, 13 billion active, with a context window of one million tokens. The model weights are available under the MIT license on Hugging Face, and the model is also available through OpenRouter.
According to DeepSeek, the model has improved in all tested categories compared with the previous version, most notably in agentic tasks. On the GDPval benchmark, which tests models on complex real-world office work, the score rose from 1189 to 1559 Elo points. According to the company, the model also hallucinates less often.
Independent commentator Simon Willison tested the model through OpenRouter using his „pelican on a bicycle" test. At the default reasoning level, he described the result as weak, while at a high reasoning level (reasoning effort high), it was significantly better.
Why it matters
According to the reported figures, the model offers performance almost on par with the more expensive GPT-5.6 Luna at a fraction of the price, which is relevant both to individual developers choosing an API for specific tasks and to companies with high query volumes, where the higher cache discount for repeated context also makes a difference. The open MIT license and the availability of weights on Hugging Face also allow the model to be run outside the provider API.
What was added since the original report
Verified updates
-
Specific prices: $0.14/million input, $0.27/million output (the existing article only gives percentages); Outperforms MiniMax M3 according to the Intelligence Index (an alternative benchmark versus GPT-5.6 Luna); Available through OpenRouter; Performance improves more markedly with a higher reasoning_effort level; Primary source: Simon Willison (independent analyst, reputation 88)
- Specific prices: $0.14/million input, $0.27/million output (the existing article only gives percentages)
- Outperforms MiniMax M3 according to the Intelligence Index (an alternative benchmark versus GPT-5.6 Luna)
- Available through OpenRouter
- Performance improves more markedly with a higher reasoning_effort level
- Primary source: Simon Willison (independent analyst, reputation 88)
Two audiences, two different impacts
What this means
For individuals
Anyone choosing a model for development with LLMs based on value for money gains another inexpensive option with performance close to more expensive models, but the result depends substantially on the selected reasoning level.
For a business
According to the reported figures, companies running LLMs at higher volumes can reduce costs per task by approximately 60 % compared with GPT-5.6 Luna at comparable performance, with additional savings on repeated context thanks to a cache discount of 98 % compared with the usual 90 %.
DevelopmentCheck the original
Event sources
confirmed by 2 independent sources · 2 publishers, 2 independent. We count feeds from the same owner only once.