Bargo
DeepSeek

DeepSeek V4 Flash 0731 Lands. OpenAI Had Already Cut Prices.

A pure post-training overhaul pushed DeepSeek's cheapest model within one point of GPT-5.6 Luna. But OpenAI cut Luna's price 80% the day before — the inference price war is on.

Bargo · 2026-07-31

DeepSeek released V4-Flash-0731 into public beta this morning, July 31. The numbers are startling. It is the same 284 billion parameter model as the April preview, but a pure post-training overhaul pushed its agent benchmark scores up by as much as 645 percent. At $0.14 per million input tokens and $0.28 per million output, it is the best price-performance model on the market.

But the timing matters. OpenAI had already cut GPT-5.6 Luna's price 80% the day before, to $0.20 per million input and $1.20 per million output. The inference price war is no longer coming — it is here.

The post-training leap

DeepSeek did not add parameters. It did not change the architecture. The 284B total parameter count and 13B active parameters are identical to the April preview. The only thing that changed was the post-training recipe, and the gains are concentrated entirely in agent capabilities.

The coding agent benchmark DeepSWE jumped from 7.3 to 54.4. Cybergym went from 38.7 to 76.7. Terminal Bench 2.1 rose from 61.8 to 82.7. On the GDPval-AA v2 evaluation, which measures agentic real-world work tasks, the Elo rating climbed from 1,189 to 1,559, a gain of 370 points. Artificial Analysis notes that V4 Flash 0731 is now within one Intelligence Index point of both GPT-5.6 Luna and GLM-5.2.

The model is also more token efficient. It used 206 million output tokens to run the Intelligence Index, against 234 million for the previous version — a 12 percent reduction while scoring higher.

DeepSeek confirmed that V4 Flash 0731 "keeps the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained."

The comparison with GPT-5.6 Sol is not close

The framing of "near parity with GPT-5.6 Sol" does not hold up. On the three benchmarks cited, Sol leads by real margins.

On GPQA Diamond, V4 Flash 0731 scores 91 percent against Sol's 94.1 percent. On Humanity's Last Exam, the gap is 37 percent to 47 percent. On the broader Artificial Analysis Intelligence Index, which combines nine evaluations, Sol scores 59 against Flash's 50 — a nine point gap.

Sol also costs $5.00 per million input and $30.00 per million output, unchanged. That is 36 times and 107 times more expensive. The correct comparison is Flash 0731 versus GPT-5.6 Luna, which is now the real price war.

The price war: Flash vs Luna

Here is the matchup that actually matters.

V4 Flash 0731 GPT-5.6 Luna Advantage
Input / 1M tokens $0.14 $0.20 Flash 1.4x
Output / 1M tokens $0.28 $1.20 Flash 4.3x
Cache hit input $0.0028 Flash 98% discount
Intelligence Index 50 51 Luna +1 pt
GDPval-AA v2 Elo 1,559
Context window 1M 1.1M Parity
Open weights MIT license Closed Flash

Flash is cheaper on every pricing dimension, especially output where the gap is 4.3x. Luna is one point ahead on intelligence. The two models are in the same tier on both cost and capability — that is the story.

The open-source inference engine

The weights are on HuggingFace under the MIT license. Unsloth GGUFs are already available, meaning a single RTX 5090 can run the model at home. The 98 percent cache hit discount on the first-party API drops the effective input price to $0.0028 per million tokens — the most aggressive discount in the industry.

The provider landscape is still catching up. OpenRouter has the 0731 version live. DeepSeek's own API automatically routes the model name deepseek-v4-flash to the new version. Together AI and Fireworks are still serving the older V4 Flash. That will likely change once the full weight files propagate, which Artificial Analysis expects in the coming weeks.

What comes next

DeepSeek confirmed that V4-Pro is "coming ASAP." The V4-Pro API and web and app models are unchanged today. Only Flash received the upgrade. The Responses API format, which Flash now supports natively, is adapted for Codex, and Pro will gain that support in early August.

The inference price war is now a two-sided fight. OpenAI cut Luna 80 percent on July 30. DeepSeek shipped a near-Luna model at Luna-beating prices on July 31. The gap between the two is now measured in single-digit percentages on both price and performance. The winner is anyone building on top of these models — the cost of intelligence is in free fall.


More research at bargo.ai/research.

Get Bargo research in your inbox
One email when we publish. No spam, unsubscribe anytime.
Get Bargo research in your inbox
One email when we publish. No spam.