V4 Flash Vision Exp launch and availability

DeepSeek on Friday released an experimental multimodal large language model called V4 Flash Vision Exp, derived from its text-only V4 Flash model. The Hangzhou-based company is initially offering the new model through its paid developer API and may later publish an open-source version, following the pattern of earlier V4 releases. V4 Flash Vision Exp is the first multimodal addition to DeepSeek's V4 series, which previously contained only text-oriented models. According to Deccan Chronicle, the company described the model as matching V4 Flash on text capabilities including agents, reasoning, and world knowledge, while adding the ability to interpret images and screenshots.
Benchmark results against V4 Flash and Opus 4.8
DeepSeek compared V4 Flash Vision Exp against its predecessor across seven text-based benchmarks, with the new model outperforming V4 Flash on all but one. The exception was Cybergym, which evaluates a model's ability to find software vulnerabilities. On four visual benchmarks, V4 Flash Vision Exp scored more than 10 percent higher than V4 Flash on two. The model also bested Anthropic's Opus 4.8 on two visual tests called ALE and ZeroBench. ALE contains more than 1,000 multi-step tasks requiring models to interact with applications, write code, and interpret media files, while ZeroBench contains 100 image-analysis tasks designed to challenge frontier systems. The Next Web reported that DeepSeek's own published comparison table shows V4 Flash Vision Exp winning three of eleven head-to-head benchmarks against Opus 4.8.
Underlying V4 Flash architecture
DeepSeek has not disclosed architectural details specific to V4 Flash Vision Exp, but its Hugging Face page documents the parent V4 Flash model. V4 Flash is a mixture-of-experts model with 284 billion parameters spread across multiple neural networks of 13 billion parameters each; only the most relevant network activates per query. Two techniques called HCA and CSA compress the model's KV cache, reducing the compute needed to process prompts of 1 million tokens by 73 percent. The model was trained on 32 trillion tokens using an algorithm called Muon that accelerates calibration of hidden layers. A larger sibling, V4 Pro, carries more than five times as many parameters, raising the possibility that future specialized models could be derived from it.
Competitive context for Chinese AI developers
DeepSeek's earlier release of its flagship model this year shifted expectations for what low-cost, open-weight systems can achieve, and Chinese developers continue to compete closely with U.S. frontier labs on price-performance. The new release extends that rivalry into multimodal agent capabilities, a domain where Anthropic's Opus 4.8 has been positioned as a benchmark. DeepSeek has previously shipped multimodal and visual models, including its DeepSeek-VL family, but V4 Flash Vision Exp is the first such offering layered onto its most advanced V4 series.
Open questions and next milestones
DeepSeek has not stated whether V4 Flash Vision Exp will move from paid API access to a free tier, nor has it confirmed whether a comparable multimodal variant will be built on V4 Pro. Independent third-party evaluations of the model's multimodal agent claims have not yet been published. Subsequent reporting is expected to focus on pricing tiers, any open-weight release, and benchmark verifications against Opus 4.8 beyond DeepSeek's internal tables.
Share this article







