On August 21, DeepSeek launched V4-Flash-Vision-Exp, an experimental multimodal vision model via its API platform. No press conference. No CEO hype video. Just an API endpoint and a pricing page. The model is labeled "experimental"—not recommended for production. That label is not a disclaimer. It's a strategy.
DeepSeek Finally Has Eyes—But the Target Isn't Chat
V4-Flash-Vision-Exp is built on the V4-Flash foundation. Text capabilities—reasoning, knowledge, agent tasks—remain identical to the standard Flash model. The difference is on visual agent benchmarks, where the new model jumped from 26.2 to 36.5 on ApexBench (Pass@1), surpassing Opus 4.8's 39.4 by a narrower margin than the gap suggests.
On Agents' Last Exam, V4-Flash-Vision-Exp scored 27.3, beating Opus 4.8's 25.7. On ZeroBench (Pass@5), it scored 35.0, edging out Opus 4.8's 34.0. On Toolathlon-Verified, the gap to Opus 4.8 narrowed to just 0.3 points (75.9 vs 76.2).
But the model still has text limitations: on NL2Repo, it scored 57.7 against Opus 4.8's 69.7. On Cybergym, it lags at 75.3 vs 78.3. DeepSeek is not claiming superiority across the board. It is claiming agent capability where it matters for its strategy.

"Experimental" Is a Feature, Not a Bug
DeepSeek explicitly labels V4-Flash-Vision-Exp as "experimental"—not recommended for production use. But the model is offered through the same API as the production models, with the same pricing, and developers are already using it to build real applications.
The "experimental" tag serves two purposes: it sets expectations low while the model iterates, and it buys DeepSeek time to collect real-world feedback without committing to a product roadmap. As one analyst noted, "Experimental models are the new beta—and beta is the new launch."
The strategy is working. Developers have already shown that V4-Flash-Vision-Exp can generate full PPTs from a prompt, rebuild websites with specific visual design requirements, and create dynamic front-end demos with animated 3D elements.
Harness Gets Eyes First—The Agent Strategy
The day before the vision model launched, DeepSeek Harness v0.1.0-rc.8 dropped with native image support. The update added /goal and /plan commands that accept images, file attachments, and session references.
The timing matters. Harness is DeepSeek's open-source agent framework, designed to turn its models into autonomous workers. With vision support, Harness can now process screenshots, UI mockups, and charts in addition to text and code.
For developers, the workflow is now seamless: paste a screenshot of a bug, describe the fix, and Harness can execute it. Upload a UI design, and Harness can generate the code. The vision model is not a standalone product—it's the missing piece of the agent platform.
The "experimental" vision model also highlights a broader DeepSeek strategy: Harness is open source and model-agnostic, supporting Claude Code and Codex as sub-agents. Even if developers don't use DeepSeek models, they can use DeepSeek's agent framework. The vision model makes that framework more capable—and more likely to lock developers into the ecosystem.
The Price Is the Strategy
V4-Flash-Vision-Exp follows V4-Flash's pricing: 1 yuan per million input tokens and 2 yuan per million output tokens. Images are converted to tokens, capped at 384 tokens per image. During peak hours, that works out to roughly 0.001 yuan per image—about $0.00014.
Compare that to Opus 4.8's $25 per million output tokens: the price difference is roughly 99%. For developers building agent workflows with heavy image inputs, the economics shift dramatically. DeepSeek's play is not just performance—it's total cost of operation.
The launch also includes a free Files API, allowing developers to upload images once and reference them via file_id across multiple requests. For agent applications that repeatedly process the same images, this can further reduce costs.

What It Means
DeepSeek's vision model is not a ChatGPT competitor. It's an agent infrastructure play. The "experimental" label, the Harness integration, the price—all point to a company that sees the future of AI not as conversation, but as execution. The question isn't whether DeepSeek can build a better chatbot. It's whether it can build the operating system for AI agents. And with V4-Flash-Vision-Exp, it just added a crucial layer.
P.S. The vision model is already live. Developers can try it with the API endpoint deepseek-v4-flash-vision-exp. The model is experimental, but the pricing is final. The window to test the cheapest multimodal agent model at scale is open. And DeepSeek is already collecting the data that will make it better.
Frequently Asked Questions
Q: What is DeepSeek V4-Flash-Vision-Exp?
A: It is DeepSeek's first multimodal vision model, launched on August 21, 2026. It's an "experimental" model that adds visual understanding capabilities to the V4-Flash foundation, while maintaining the same text capabilities and pricing.
Q: How does the vision model compare to Claude Opus 4.8?
A: On visual agent benchmarks, V4-Flash-Vision-Exp matches or slightly exceeds Opus 4.8—scoring 27.3 vs 25.7 on Agents' Last Exam, and 35.0 vs 34.0 on ZeroBench (Pass@5). On text benchmarks, it still lags behind Opus 4.8 in several categories.
Q: What does "experimental" mean for this model?
A: DeepSeek labels V4-Flash-Vision-Exp as "experimental" and does not recommend it for production use. However, the model is offered through the same API with the same pricing as production models. The label signals that the model is still iterating and collecting feedback.
Q: How does DeepSeek's pricing compare to competitors for vision tasks?
A: Images are capped at 384 tokens per image. At peak pricing, that's roughly 0.001 yuan ($0.00014) per image—about 1/99th the cost of Opus 4.8. For developers building agent workflows with heavy image inputs, the economics shift dramatically.
Q: What is the relationship between the vision model and DeepSeek Harness?
A: DeepSeek Harness, the company's open-source agent framework, got native image support the day before the vision model launched. The two releases are designed to work together—Harness can now process screenshots, UI mockups, and charts in addition to text and code. This positions DeepSeek as an agent infrastructure provider, not just a model provider.
Q: What can developers build with V4-Flash-Vision-Exp?
A: Developers are already using it to generate full PPTs from prompts, rebuild websites with specific visual design requirements, create dynamic front-end demos with animated 3D elements, and build agent workflows that process screenshots and UI mockups.
Q: How are images priced in the model?
A: Images are automatically scaled to 800×800 pixels. Each image is converted to tokens with a cap of 384 tokens per image. The pricing follows V4-Flash's standard rate: 1 yuan per million input tokens and 2 yuan per million output tokens.
Q: Is DeepSeek planning to release the model weights?
A: DeepSeek has not announced plans to open-weight this model. However, DeepSeek has previously released early models as open weight, and industry observers note that a free version "is not impossible" given the company's history of open sourcing.
Q: What is the Files API?
A: The Files API allows developers to upload images once and reference them via file_id across multiple requests. For agent applications that repeatedly process the same images, this can further reduce costs. The Files API is currently free.
Q: What is the significance of this launch for the AI industry?
A: DeepSeek is positioning itself as an agent infrastructure provider—not just a model provider. The vision model is not a ChatGPT competitor; it's the missing piece of an agent framework. The "experimental" label, Harness integration, and aggressive pricing all point to a company that sees the future of AI as execution, not conversation.
