Skip to content
important New models

DeepSeek has released V4 Flash 0731 with performance close to GPT-5.6 Luna at a significantly lower price

confirmed by 2 independent sources updated August 1, 2026

DeepSeek has released V4 Flash 0731 with a score of 50 points in the Artificial Analysis Intelligence Index, one point below GPT-5.6 Luna, but at a price about 60 % lower. The weights are under the MIT license on Hugging Face, with a price of 0.14 $/0.27 $ per million tokens.

DeepSeek has released a new V4 Flash model designated „0731", an update to its affordable model. According to the Artificial Analysis Intelligence Index, the model scores 50 points, ten more than the previous V4 Flash version from April 2026. The score is just one point lower than that of GPT-5.6 Luna from OpenAI, yet according to the sources, a task with V4 Flash is approximately 60 percent cheaper than with GPT-5.6 Luna, even after its 80 percent price reduction. According to Artificial Analysis, the model also outperforms MiniMax M3 (a model with 428 billion parameters) in the Intelligence Index.

The price is 0.14 dollars per million input tokens and 0.27 dollars per million output tokens. DeepSeek reports a cache discount of 98 percent, above the industry standard of 90 percent. According to the company, the model uses 12 percent fewer tokens than the previous version. The architecture remains the same: 284 billion parameters in total, 13 billion active, with a context window of one million tokens. The model weights are available under the MIT license on Hugging Face, and the model is also available through OpenRouter.

According to DeepSeek, the model has improved in all tested categories compared with the previous version, most notably in agentic tasks. On the GDPval benchmark, which tests models on complex real-world office work, the score rose from 1189 to 1559 Elo points. According to the company, the model also hallucinates less often.

Independent commentator Simon Willison tested the model through OpenRouter using his „pelican on a bicycle" test. At the default reasoning level, he described the result as weak, while at a high reasoning level (reasoning effort high), it was significantly better.

What changed

Why it matters

According to the reported figures, the model offers performance almost on par with the more expensive GPT-5.6 Luna at a fraction of the price, which is relevant both to individual developers choosing an API for specific tasks and to companies with high query volumes, where the higher cache discount for repeated context also makes a difference. The open MIT license and the availability of weights on Hugging Face also allow the model to be run outside the provider API.

What was added since the original report

Verified updates

  1. New verified information

    Specific prices: $0.14/million input, $0.27/million output (the existing article only gives percentages); Outperforms MiniMax M3 according to the Intelligence Index (an alternative benchmark versus GPT-5.6 Luna); Available through OpenRouter; Performance improves more markedly with a higher reasoning_effort level; Primary source: Simon Willison (independent analyst, reputation 88)

    • Specific prices: $0.14/million input, $0.27/million output (the existing article only gives percentages)
    • Outperforms MiniMax M3 according to the Intelligence Index (an alternative benchmark versus GPT-5.6 Luna)
    • Available through OpenRouter
    • Performance improves more markedly with a higher reasoning_effort level
    • Primary source: Simon Willison (independent analyst, reputation 88)

Two audiences, two different impacts

What this means

01

For individuals

Anyone choosing a model for development with LLMs based on value for money gains another inexpensive option with performance close to more expensive models, but the result depends substantially on the selected reasoning level.

What to do Try DeepSeek V4 Flash 0731 through OpenRouter and set a higher reasoning level (reasoning effort high) for more demanding tasks, because the default setting produced weaker results according to the test.
More practical updates →
02

For a business

According to the reported figures, companies running LLMs at higher volumes can reduce costs per task by approximately 60 % compared with GPT-5.6 Luna at comparable performance, with additional savings on repeated context thanks to a cache discount of 98 % compared with the usual 90 %.

Development
What to decide Consider deploying DeepSeek V4 Flash 0731 as a cheaper alternative for operations with a high volume of queries, especially where the same context is reused repeatedly (cache discount 98 %).
More business impacts →
API benchmarking price price comparison DeepSeek GPT 5.6-Luna LLM model OpenRouter

Check the original

Event sources

confirmed by 2 independent sources · 2 publishers, 2 independent. We count feeds from the same owner only once.

2
The Decoder (daily AI news) independent context · first detected New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost Simon Willison — AI tag (leading independent LLM commentator) community signal deepseek-ai/DeepSeek-V4-Flash-0731