Models

DeepSeek V4.1 Flash Is Here—40 Days After V4 Pro, and It Might Replace It

CRAZE CRAZE Summary 3 things to know
  • DeepSeek V4.1 Flash is in internal testing—new native multimodal architecture, 3x faster than V4 Flash.
  • The model may replace V4 Pro entirely just 40 days after Pro launched, according to the company's internal survey.
  • V4 Pro requests will be auto-routed to Flash at Flash pricing—an industry-rare "upgrade and discount" package.
Jeff Editorial | · 4 min read
DeepSeek V4.1 Flash Is Here—40 Days After V4 Pro, and It Might Replace It

On September 8, DeepSeek opened a limited internal test of V4.1 Flash, accessible via the endpoint deepseek-v4.1-flash-expires-on-0910. The endpoint expires on September 10, but the company has confirmed the model is on track for official release around that date.

The announcement was quiet—no press release, no technical report, just an API endpoint and an internal questionnaire. But the implications are anything but quiet.

New Architecture, Native Multimodality—No Bolt-On

The most significant change is architectural. Previous V4 Flash Vision-Exp was a "bolt-on"—a vision encoder and aligner attached to a text-only base model. V4.1 Flash is designed from the ground up with native multimodal support, accepting both image and text inputs in a single unified architecture.

This is the first time DeepSeek has integrated visual understanding into the model itself, rather than as an add-on. The move puts V4.1 Flash on par with how OpenAI and Google structure their flagship models.

3x Faster. Same Price. Pro-Level Answers.

Developers testing the intermediate checkpoint reported generation speeds averaging 284 tokens per second, nearly three times the V4 Flash baseline of about 97 tokens per second. Some tasks hit 420 tokens per second. One developer's "hello" prompt returned in 0.8 seconds end-to-end.

The speed gains come from new architecture—not just post-training. In 49K-token retrieval tasks, V4.1 Flash was 5.2x faster; in SVG code generation, 6.0x faster.

Pricing remains at V4 Flash levels: 4 yuan per million output tokens (idle) and 8 yuan per million (peak), with cache-hits dropping to 0.02 yuan on September 10. The price of V4 Flash will drop further upon the model's full release.

But speed doesn't always mean lower cost. One early tester ran 3 billion tokens across 14 tasks and found that while V4.1 Flash generated faster, it used more output tokens on complex tasks and consumed more total compute, increasing cost by about 36% in that specific test. The trade-off is being documented in real time.

The Survey Question: "Can This Replace Pro?"

The most telling detail is the feedback survey. One question asks developers directly: "Do you think this model can fully replace the live DeepSeek V4 Pro?"

This is not a hypothetical. V4 Pro launched just 40 days ago. If the answer is yes, DeepSeek may effectively deprecate its own flagship with a Flash-tier model. The company has confirmed that upon V4.1 Flash's full release, V4 Pro requests will be automatically routed to V4.1 Flash and billed at Flash rates—an "upgrade and discount" package that is rare in the industry.

DeepSeek V4.1 Flash Is Here—40 Days After V4 Pro, and It Might Replace It
DeepSeek V4.1 Flash uses native multimodality, runs 3x faster, and may replace V4 Pro entirely—just 40 days after Pro launched.

DeepSeek just demonstrated that it can iterate from V4 Pro to V4.1 Flash in 40 days—with new architecture, native multimodality, and performance that beats its own flagship. The speed of iteration is the story.

The company is running a live test of whether a cheaper, faster model can replace its premium offering. If developers and internal benchmarks confirm it, V4 Pro may be quietly retired. That's not a product roadmap. That's a product rethink.


P.S. The "expires-on-0910" endpoint is temporary. But the strategy behind it is permanent: DeepSeek is compressing its release cycles, blurring the line between tiers, and using early developer feedback to shape the next model before the current one finishes its first month. The question isn't whether V4.1 Flash can replace V4 Pro. It's whether V4.1 Pro will replace V4.1 Flash before the year ends.


Frequently Asked Questions

Q: What is DeepSeek V4.1 Flash?

A: DeepSeek V4.1 Flash is an intermediate internal test model that began rolling out on September 8, 2026. It uses a new architecture with native multimodal support and is positioned as a potential replacement for V4 Pro.

Q: How fast is V4.1 Flash?

A: Developers reported generation speeds averaging 284 tokens per second, nearly 3x faster than V4 Flash's baseline of about 97 tokens per second. Some tasks hit 420 tokens per second.

Q: Is V4.1 Flash replacing V4 Pro?

A: DeepSeek is asking developers in a feedback survey: "Do you think this model can fully replace the live V4 Pro?" The company has confirmed that upon V4.1 Flash's full release, V4 Pro requests may be automatically routed to V4.1 Flash and billed at Flash rates.

Q: When will V4.1 Flash be fully released?

A: The test endpoint expires on September 10, 2026, and the company has confirmed the model is on track for official release around that date.

Q: What's new in V4.1 Flash?

A: Native multimodal support—the first DeepSeek model with vision built into the base architecture, rather than an external vision encoder attached to a text model.

Q: Does faster generation mean lower cost?

A: Not necessarily. In one early test across 14 tasks using 3 billion tokens, V4.1 Flash consumed about 36% more total compute due to generating longer outputs, despite faster generation speed.

Q: How much does V4.1 Flash cost?

A: Pricing remains at V4 Flash levels: 4 yuan per million output tokens (idle) and 8 yuan per million (peak), with cache-hits dropping to 0.02 yuan on September 10. V4 Flash pricing is also expected to drop further.

Q: How long after V4 Pro did V4.1 Flash appear?

A: V4 Pro launched on August 13, 2026—just 40 days before V4.1 Flash began testing.

Q: Why is DeepSeek moving so fast?

A: DeepSeek is compressing its release cycles and using early developer feedback to shape the next model before the current one finishes its first month. The company is blurring the line between product tiers.

Q: Is V4.1 Flash the same as V4 Flash Vision?

A: No. V4 Flash Vision-Exp was a "bolt-on"—a vision encoder attached to a text model. V4.1 Flash is a new architecture with native multimodality built from the ground up.

Advertisement

CRAZE

Use CRAZE to turn this article into a faster answer: pull the summary, surface the key term, or jump straight to the next story in this thread.

Article