DeepSeek V4 GA is here. 1.6 trillion parameters. 1 million token context. Pro version at $0.87 per million output tokens off-peak. Flash version at $0.28.
It is not the most powerful model on the market. But it might be the best value.
The official version launched on July 20, following a preview release in April. It comes in two variants: DeepSeek-V4-Pro (1.6T total parameters, 49B activated) and DeepSeek-V4-Flash (284B total parameters, 13B activated).

The numbers are clear. On LiveCodeBench, V4-Pro scores 93.5 percent. Codeforces Elo: 3,206. On agentic coding benchmarks, V4-Pro has reached best-in-class open-source levels, with internal feedback indicating it outperforms Sonnet 4.5 and approaches Opus 4.6 in non-thinking mode. Agent capabilities are significantly stronger, with major gains in 3D and SVG generation. Developer Pankaj Kumar summed it up directly: overall performance is close to Opus 4.8 level, coding approaches GPT-5.6 Sol.
But the real story is pricing.
Model | Price per 1M output tokens | Performance tier |
|---|---|---|
DeepSeek V4 Pro | $0.87 (off-peak) | Opus 4.8-class |
DeepSeek V4 Flash | $0.28 (off-peak) | Near Pro |
Claude Fable 5 | $50 | Flagship |
GPT-5.6 Sol | $30 | Flagship |
Kimi K3 | $15 | Open-weight challenger |
V4 Pro costs 1/57 of Fable 5. 1/34 of GPT-5.6 Sol. 1/17 of Kimi K3.
DeepSeek also introduced peak-hour pricing for the first time. Peak hours: 9:00-12:00 and 14:00-18:00 Beijing time, with prices doubling. Pro output: $0.87 off-peak, $1.74 peak. Flash output: $0.28 off-peak, $0.56 peak. Cache-hit input pricing drops as low as $0.00025 per million tokens. AI compute is now priced like electricity.
Industry analysts see this as a sign of AI cloud services maturing. Peak-hour pricing will reshape developer habits, pushing batch workloads to off-peak hours and making task orchestration a new competitive advantage for cost control.
The older models — deepseek-chat and deepseek-reasoner — will be deprecated on July 24. All API calls, workflows, and agent applications must migrate within four days.

V4 is also DeepSeek's first flagship model optimized for Huawei's Ascend chips. Huawei has announced full support for DeepSeek V4 across its Ascend supernode product line, with model release and compute adaptation synchronized. Nvidia CEO Jensen Huang previously warned: "A new DeepSeek model built on Huawei's platform would be a bad outcome for the US."
V4 uses a new attention mechanism that drastically reduces long-context compute costs. At 1M context, V4-Pro uses only 27 percent of V3.2's compute and 10 percent of its memory. V4-Flash is even more efficient: 10 percent compute, 7 percent memory. The 1M context window is now standard across all official services, not a premium add-on.
P.S. If you are a Fable 5 user, V4 Pro costs 1/57 of what you are paying, with performance close enough that you might not notice the difference. If you are a developer, the clock is ticking — July 24 is the migration deadline. If you are a workload scheduler, it is time to set an alarm for 3 AM. That is when the tokens are cheapest.
