Models

DeepSeek Just Turned AI Compute Into Electricity. Peak Pricing Is Here.

CRAZE CRAZE Summary 3 things to know
  • DeepSeek's new peak/off-peak pricing charges double during weekday business hours, with V4-Pro output up 350% and cached input up 12x.
  • The era of subsidized AI APIs is ending as DeepSeek and Chinese rivals shift from price wars to sustainable economics under compute constraints.
  • China's domestic chipmakers like Huawei and Cambricon emerge as real winners, as constrained GPU supply opens a pricing window amid soaring token demand.
Jeff Editorial | · 5 min read
DeepSeek Just Turned AI Compute Into Electricity. Peak Pricing Is Here.

On August 13, DeepSeek announced new API pricing with a peak/off-peak structure: peak hours (9:00-12:00 and 14:00-18:00 Beijing time) cost double the off-peak rate. The new prices take effect August 17.

The numbers are striking. V4-Pro output tokens currently cost 6 yuan per million. After August 17, off-peak output jumps to 13.5 yuan; peak output hits 27 yuan—a 350% increase. For cached input tokens, the peak price rises from 0.025 yuan to 0.30 yuan per million: a 12x increase.

This was never a mystery. DeepSeek had been warning users for months that prices would rise "significantly" once V4-Pro exited preview. But the structure—not just the magnitude—is what matters.

MODEL

deepseek-v4-flash

deepseek-v4-pro

BASE URL (OpenAI Format)

https://api.deepseek.com

BASE URL (Anthropic Format)

https://api.deepseek.com/anthropic

MODEL VERSION

DeepSeek-V4-Flash-0731

DeepSeek-V4-Pro-0813

THINKING MODE

Supports both non-thinking and thinking (default) modes
See
Thinking Mode for how to switch

CONTEXT LENGTH

1M

MAX OUTPUT

MAXIMUM: 384K

FEATURES

Json Output

Tool Calls

Responses API

Anthropic API

Chat Prefix Completion(Beta)

FIM Completion(Beta)

Non-thinking mode only

Non-thinking mode only

PRICING(1)

1M INPUT TOKENS (CACHE HIT)

$0.0028

$0.003625

1M INPUT TOKENS (CACHE MISS)

$0.14

$0.435

1M OUTPUT TOKENS

$0.28

$0.87

Concurrency Limit(2)

2500

500

12x on Input, 350% on Output—The Math Hits Hard

The peak/off-peak model is borrowed from electricity grids. Daytime business hours see concentrated AI workloads—model calls, intelligent interactions, data processing—driving GPU cluster loads to capacity. Night and weekends see expensive hardware idling.

DeepSeek is using price to solve a physical problem: compute cannot be stored. The only way to balance supply and demand is to shift demand to where supply is abundant. Peak pricing is the market signal that says: "Move your non-urgent work to off-peak hours, or pay for the privilege of real-time access."

This is not a pricing innovation. It is a supply constraint.

MODEL

1M INPUT TOKENS (CACHE HIT)

1M INPUT TOKENS (CACHE MISS)

1M OUTPUT TOKENS

deepseek-v4-flash

OFF-PEAK

$0.007

$0.22

$0.66

PEAK

$0.014

$0.44

$1.32

deepseek-v4-pro

OFF-PEAK

$0.022

$0.66

$1.98

PEAK

$0.044

$1.32

$3.96

"Ferrari at Scooter Prices" Was Never Sustainable

For the past year, DeepSeek has been the industry's pricing disrupter. V4-Flash became the most-used model on OpenRouter's aggregation platform—not because it was the best, but because it was the cheapest. "Ferrari performance at scooter prices" was the strategy that bought market share.

But market share at a loss is not a business model. As one developer put it after the announcement: "It's no longer cheap. The price is now on par with similar models." The era of subsidized API calls is ending.

DeepSeek is not alone. Zhipu has raised API prices three times this year. Moonshot's Kimi K3 pricing is 3-4x higher than its predecessor. Tencent Cloud, Alibaba Cloud, and Baidu Cloud have all raised AI compute prices. What was once a price war is now a coordinated retreat from margin destruction.

The Real Winner Isn't DeepSeek—It's China's Chip Makers

The most important consequence of DeepSeek's price hike is not about API pricing. It is about what the price hike reveals.

Demand for inference tokens is surging. Alibaba's token usage has grown 1,000x since its 2024 launch. ByteDance spent roughly 90 billion yuan on AI compute in 2025. Tencent's first-half capex reached 84.7 billion yuan—double its full-year 2025 total. But supply is constrained.

That supply constraint is the opening for China's AI chip industry. Huawei's Ascend shipped about 812,000 units in 2025; Cambricon shipped about 116,000. Nvidia's share of the Chinese market has fallen from roughly 95% to about 55%. As Huatai Securities put it: the combination of rising costs and domestic supply-demand imbalance is opening a "pricing window" for domestic AI chips.

The API price hike is bad news for developers. It is good news for the companies that make the chips running those models. And as the price of compute rises, the value of domestic chip capacity rises with it.

DeepSeek Just Turned AI Compute Into Electricity. Peak Pricing Is Here.
DeepSeek's peak pricing treats AI compute like electricity—more expensive during the day, cheaper at night.

What It Means

DeepSeek's peak pricing is not a one-time adjustment. It is a preview of how AI infrastructure will be priced in a compute-constrained world. For developers, the free ride is ending. For the industry, the shift from price war to sustainable economics has begun.


P.S. The actual deadline is August 17. Developers who want to run V4-Pro at current prices have until midnight Beijing time. After that, the cheap API calls are gone—and the industry moves into a new phase where compute is no longer treated as infinite.


Frequently Asked Questions

Q: When does DeepSeek's new peak pricing take effect?

A: August 17, 2026.

Q: What are the peak hours?

A: Two weekday periods: 9:00-12:00 and 14:00-18:00 Beijing time. Off-peak covers all other hours, including weekends.

Q: How much are prices increasing?

A: V4-Pro output tokens: from 6 yuan to 27 yuan per million during peak hours (350% increase). Cached input tokens: from 0.025 yuan to 0.30 yuan per million during peak hours (12x increase).

Q: Is DeepSeek the only company raising prices?

A: No. Zhipu has raised API prices three times this year. Moonshot's Kimi K3 pricing is 3-4x higher than its predecessor. Tencent Cloud, Alibaba Cloud, and Baidu Cloud have all raised AI compute prices.

Q: Why is DeepSeek adopting peak/off-peak pricing?

A: Compute cannot be stored. Peak pricing is a market signal to shift non-urgent work to off-peak hours when GPU capacity is abundant. It reflects the physical reality that during business hours, compute clusters are at capacity, while they idle at night and on weekends.

Q: Does this affect all DeepSeek models?

A: The price change applies to V4-Pro and V4-Flash. Prices for V3 models and earlier are not affected at this time.

Q: How does the new pricing compare to competitors?

A: Even after the increase, DeepSeek's peak pricing remains roughly on par with or below comparable models from Zhipu, Moonshot, and other domestic providers. But the "Ferrari at scooter prices" era is over.

Q: What does this mean for developers?

A: Developers who can move batch processing to off-peak hours will see minimal impact. Real-time, latency-sensitive applications running during peak hours will face significantly higher costs.

Q: Is DeepSeek losing money on API calls?

A: The company has not disclosed financials, but the broad industry shift from price wars to sustainable pricing suggests that the previous pricing model was not sustainable for any provider.

Q: Who benefits from this price hike?

A: The most direct beneficiaries are domestic chip makers—Huawei Ascend, Cambricon, and others. Rising API prices signal compute scarcity, which increases the value of domestic AI chip capacity. As one analyst put it: the price hike is a "pricing window" for domestic chips.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article