On August 13, DeepSeek announced new API pricing with a peak/off-peak structure: peak hours (9:00-12:00 and 14:00-18:00 Beijing time) cost double the off-peak rate. The new prices take effect August 17.
The numbers are striking. V4-Pro output tokens currently cost 6 yuan per million. After August 17, off-peak output jumps to 13.5 yuan; peak output hits 27 yuan—a 350% increase. For cached input tokens, the peak price rises from 0.025 yuan to 0.30 yuan per million: a 12x increase.
This was never a mystery. DeepSeek had been warning users for months that prices would rise "significantly" once V4-Pro exited preview. But the structure—not just the magnitude—is what matters.
MODEL | deepseek-v4-flash | deepseek-v4-pro | |
BASE URL (OpenAI Format) | |||
BASE URL (Anthropic Format) | |||
MODEL VERSION | DeepSeek-V4-Flash-0731 | DeepSeek-V4-Pro-0813 | |
THINKING MODE | Supports both non-thinking and thinking (default) modes | ||
CONTEXT LENGTH | 1M | ||
MAX OUTPUT | MAXIMUM: 384K | ||
FEATURES | ✓ | ✓ | |
✓ | ✓ | ||
✓ | ✓ | ||
✓ | ✓ | ||
✓ | ✓ | ||
Non-thinking mode only | Non-thinking mode only | ||
PRICING(1) | 1M INPUT TOKENS (CACHE HIT) | $0.0028 | $0.003625 |
1M INPUT TOKENS (CACHE MISS) | $0.14 | $0.435 | |
1M OUTPUT TOKENS | $0.28 | $0.87 | |
Concurrency Limit(2) | 2500 | 500 | |
12x on Input, 350% on Output—The Math Hits Hard
The peak/off-peak model is borrowed from electricity grids. Daytime business hours see concentrated AI workloads—model calls, intelligent interactions, data processing—driving GPU cluster loads to capacity. Night and weekends see expensive hardware idling.
DeepSeek is using price to solve a physical problem: compute cannot be stored. The only way to balance supply and demand is to shift demand to where supply is abundant. Peak pricing is the market signal that says: "Move your non-urgent work to off-peak hours, or pay for the privilege of real-time access."
This is not a pricing innovation. It is a supply constraint.
MODEL | 1M INPUT TOKENS (CACHE HIT) | 1M INPUT TOKENS (CACHE MISS) | 1M OUTPUT TOKENS | |
deepseek-v4-flash | OFF-PEAK | $0.007 | $0.22 | $0.66 |
PEAK | $0.014 | $0.44 | $1.32 | |
deepseek-v4-pro | OFF-PEAK | $0.022 | $0.66 | $1.98 |
PEAK | $0.044 | $1.32 | $3.96 | |
"Ferrari at Scooter Prices" Was Never Sustainable
For the past year, DeepSeek has been the industry's pricing disrupter. V4-Flash became the most-used model on OpenRouter's aggregation platform—not because it was the best, but because it was the cheapest. "Ferrari performance at scooter prices" was the strategy that bought market share.
But market share at a loss is not a business model. As one developer put it after the announcement: "It's no longer cheap. The price is now on par with similar models." The era of subsidized API calls is ending.
DeepSeek is not alone. Zhipu has raised API prices three times this year. Moonshot's Kimi K3 pricing is 3-4x higher than its predecessor. Tencent Cloud, Alibaba Cloud, and Baidu Cloud have all raised AI compute prices. What was once a price war is now a coordinated retreat from margin destruction.
The Real Winner Isn't DeepSeek—It's China's Chip Makers
The most important consequence of DeepSeek's price hike is not about API pricing. It is about what the price hike reveals.
Demand for inference tokens is surging. Alibaba's token usage has grown 1,000x since its 2024 launch. ByteDance spent roughly 90 billion yuan on AI compute in 2025. Tencent's first-half capex reached 84.7 billion yuan—double its full-year 2025 total. But supply is constrained.
That supply constraint is the opening for China's AI chip industry. Huawei's Ascend shipped about 812,000 units in 2025; Cambricon shipped about 116,000. Nvidia's share of the Chinese market has fallen from roughly 95% to about 55%. As Huatai Securities put it: the combination of rising costs and domestic supply-demand imbalance is opening a "pricing window" for domestic AI chips.
The API price hike is bad news for developers. It is good news for the companies that make the chips running those models. And as the price of compute rises, the value of domestic chip capacity rises with it.

What It Means
DeepSeek's peak pricing is not a one-time adjustment. It is a preview of how AI infrastructure will be priced in a compute-constrained world. For developers, the free ride is ending. For the industry, the shift from price war to sustainable economics has begun.
P.S. The actual deadline is August 17. Developers who want to run V4-Pro at current prices have until midnight Beijing time. After that, the cheap API calls are gone—and the industry moves into a new phase where compute is no longer treated as infinite.
Frequently Asked Questions
Q: When does DeepSeek's new peak pricing take effect?
A: August 17, 2026.
Q: What are the peak hours?
A: Two weekday periods: 9:00-12:00 and 14:00-18:00 Beijing time. Off-peak covers all other hours, including weekends.
Q: How much are prices increasing?
A: V4-Pro output tokens: from 6 yuan to 27 yuan per million during peak hours (350% increase). Cached input tokens: from 0.025 yuan to 0.30 yuan per million during peak hours (12x increase).
Q: Is DeepSeek the only company raising prices?
A: No. Zhipu has raised API prices three times this year. Moonshot's Kimi K3 pricing is 3-4x higher than its predecessor. Tencent Cloud, Alibaba Cloud, and Baidu Cloud have all raised AI compute prices.
Q: Why is DeepSeek adopting peak/off-peak pricing?
A: Compute cannot be stored. Peak pricing is a market signal to shift non-urgent work to off-peak hours when GPU capacity is abundant. It reflects the physical reality that during business hours, compute clusters are at capacity, while they idle at night and on weekends.
Q: Does this affect all DeepSeek models?
A: The price change applies to V4-Pro and V4-Flash. Prices for V3 models and earlier are not affected at this time.
Q: How does the new pricing compare to competitors?
A: Even after the increase, DeepSeek's peak pricing remains roughly on par with or below comparable models from Zhipu, Moonshot, and other domestic providers. But the "Ferrari at scooter prices" era is over.
Q: What does this mean for developers?
A: Developers who can move batch processing to off-peak hours will see minimal impact. Real-time, latency-sensitive applications running during peak hours will face significantly higher costs.
Q: Is DeepSeek losing money on API calls?
A: The company has not disclosed financials, but the broad industry shift from price wars to sustainable pricing suggests that the previous pricing model was not sustainable for any provider.
Q: Who benefits from this price hike?
A: The most direct beneficiaries are domestic chip makers—Huawei Ascend, Cambricon, and others. Rising API prices signal compute scarcity, which increases the value of domestic AI chip capacity. As one analyst put it: the price hike is a "pricing window" for domestic chips.
