On September 10, DeepSeek ended the internal test of V4.1 Flash and made the model generally available. The company confirmed that V4 Pro will be retired on September 14, with all requests automatically routed to V4.1 Flash and billed at the lower Flash rate. V4 Pro launched on August 13—its operational life was roughly one month.
The model is not a small update. V4.1 Flash uses a new architecture with native multimodal support, accepts both image and text inputs in a single unified model, and runs on the same API endpoint. It is, in every measurable dimension, the successor to V4 Pro—at a lower price.
The "Pro" Tier Just Disappeared
DeepSeek's product hierarchy has been collapsing for weeks. V4 Flash launched July 31. V4 Pro followed on August 13. V4.1 Flash arrived September 8 as an internal test. Now, on September 10, V4 Pro is being retired.
The company's own survey question—"Do you think this model can fully replace the live DeepSeek V4 Pro?"—was answered by the pricing page before users could respond. V4 Pro requests will be auto-routed to V4.1 Flash, billed at Flash rates, with no Pro premium.
The "Pro" tier no longer means "more capable." It means "more expensive for the same capability." DeepSeek just removed the premium without removing the model.
Cache Hits Cost 60% Less. Agents Just Got Cheaper to Run.
The pricing change matters more than the model itself. Under the new Flash pricing, cache-hit input tokens drop from 0.05 yuan to 0.02 yuan per million—a 60% cut. Cache-miss input falls from 1.5 yuan to 1 yuan. Output falls from 4.5 yuan to 4 yuan. Peak-hour pricing remains double the idle rate.
For agent workflows, cache hits are the dominant cost. An agent re-reading a repository, re-processing a document, or re-querying context pays for cache-hit tokens repeatedly. Cutting cache-hit prices 60% directly reduces the cost of running agents—not just the cost of running a single prompt.
DeepSeek's timing is deliberate. The model it released today is not just cheaper to call. It is cheaper to call repeatedly, which is what agents do.
3-5x Faster—and Your Balance Will Notice
Developer benchmarks from the internal test showed generation speeds of 280 to 507 tokens per second, compared to roughly 120 to 140 tokens per second for V4 Flash. In long-context retrieval and code-generation tasks, end-to-end speed improved 5-6x. One developer completed a 2D animation task in 75.4 seconds at 328 tokens per second.
But speed cuts both ways. At the same per-token price, faster generation means the balance depletes faster. If you're billed per token and the model generates faster, you run out of budget sooner. The cost efficiency gain depends on whether the speed improvement also reduces the number of tokens needed—and early testers found that V4.1 Flash sometimes generates more output on complex tasks, consuming more total compute.
The trade-off is real. DeepSeek is betting that faster, cheaper, and multimodal is worth the token-count uncertainty. For agent users, the cache-hit cut may outweigh the output increase.

DeepSeek just demonstrated that its product tiers are not fixed. The Pro model lasted 40 days. The Flash model replaced it at a lower price. The new pricing cut cache costs by 60%. This is not a product roadmap. It's a product rethink—executed in six weeks.
The question isn't whether V4.1 Flash can replace V4 Pro. It already has. The question is whether V4.1 Pro will replace V4.1 Flash before the year ends.
P.S. V4 Pro requests will be auto-routed to V4.1 Flash on September 14. If you had a V4 Pro integration, it will keep working—at the Flash price. DeepSeek didn't announce a deprecation; it just changed the price and let the model disappear. That's a quieter way to retire a flagship than most companies would choose.
Frequently Asked Questions
Q: What did DeepSeek announce on September 10?
A: DeepSeek officially released V4.1 Flash and confirmed that V4 Pro will be retired on September 14. All V4 Pro requests will be automatically routed to V4.1 Flash and billed at Flash prices.
Q: How long did V4 Pro last?
A: V4 Pro launched on August 13, 2026. It will be retired on September 14, 2026—roughly one month.
Q: What is V4.1 Flash?
A: V4.1 Flash is DeepSeek's new model with a new architecture, native multimodal support (accepting both image and text inputs), and generation speeds of 280-507 tokens per second—roughly 3-5x faster than V4 Flash.
Q: What are the new prices?
A: Cache-hit input: 0.02 yuan per million tokens (down from 0.05, a 60% cut). Cache-miss input: 1 yuan (down from 1.5). Output: 4 yuan (down from 4.5). Peak-hour pricing is double the idle rate.
Q: Why does the cache-hit price cut matter?
A: Agent workflows repeatedly process the same context, documents, or repositories. Cache-hit tokens are the dominant cost for agents. A 60% cut directly reduces the cost of running agents, not just single prompts.
Q: Is V4.1 Flash faster than V4 Pro?
A: Yes. Developer benchmarks show 280-507 tokens per second versus roughly 120-140 for V4 Flash. In long-context retrieval and code-generation tasks, end-to-end speed improved 5-6x.
Q: Does faster generation mean lower cost?
A: Not necessarily. At the same per-token price, faster generation means the balance depletes faster. Early testers found that V4.1 Flash sometimes generates more output on complex tasks, consuming more total compute.
Q: What does the "Pro" tier mean now?
A: The "Pro" tier no longer means "more capable." It means "more expensive for the same capability." DeepSeek removed the premium without removing the model.
Q: Will existing V4 Pro integrations break?
A: No. V4 Pro requests will be auto-routed to V4.1 Flash on September 14. Existing integrations will keep working—at the Flash price.
Q: Why did DeepSeek retire V4 Pro so quickly?
A: V4.1 Flash outperforms V4 Pro on performance, cost, speed, and total time. DeepSeek is compressing its release cycles and blurring the line between product tiers.
