DeepSeek posted a brief notice on August 6: "We plan to raise DeepSeek API pricing in the near future. The increase is expected to be significant. Please plan your usage accordingly." No numbers. No timeline. Just a warning.
The "Token price butcher" that once undercut everyone is now raising prices. The announcement came exactly when DeepSeek V4 Flash was processing 8 trillion tokens in a single day on the OpenCode platform.

The scale is difficult to process. 8 trillion tokens in one day. That's roughly the size of the entire internet compressed into daily API calls. 3 trillion of those tokens came from paying developers.
The week before, OpenRouter ranked DeepSeek V4 Flash as the most-used model globally with 7.22 trillion weekly tokens — ahead of every US model. On August 4, just days before the announcement, deepseek-v4-flash experienced capacity shortages due to "unprecedented access volume."
DeepSeek's current pricing is still dramatically cheap. V4-Flash costs 0.02 yuan ($0.003) per million input tokens on cache hit, 1 yuan ($0.14) on cache miss, and 2 yuan ($0.28) for output. V4-Pro is higher at 3 yuan ($0.42) input and 6 yuan ($0.84) output. Those prices are no longer sustainable at 8 trillion tokens per day.
DeepSeek is not alone. The entire Chinese AI industry is raising prices. Zhipu AI raised API prices three times this year, with Q1 2026 pricing up 83% from late 2025. Calling volume still grew 400%. Moonshot AI raised Kimi K3's API price to 20 yuan input and 100 yuan output per million tokens — roughly 4x the previous generation's pricing. MiniMax, StepFun, and other labs have followed. The price war that defined early 2026 is over.

Goldman Sachs called it: peak-hour pricing isn't a demand signal — it's a capacity signal. Demand is strong enough that pricing power is returning.
For developers who built on DeepSeek's low prices, the math just changed. The company's V4-Pro formal version hasn't even launched yet. The upcoming price hike may coincide with that release. The wider industry trend matters more than the specific numbers. The entire market is normalizing: smart developers will hedge accordingly, and architecture that assumes permanently cheap inference will need revisiting. For the industry, the announcement signals a broader realignment — the era of subsidized AI is ending.
P.S. If you're an AI developer who built on DeepSeek's low pricing, the math just changed — the "Token price butcher" is no longer bleeding margins to win market share, but processing enough volume to charge more and still grow. The cheap token era is ending, not because DeepSeek wants it to, but because the demand is too big to subsidize.
