Catch up on the essentials
- Chinese AI startup DeepSeek launched DeepSeek-V4.1-Flash on Thursday, September 10, 2026, calling it the smallest model in its new architecture family.
- V4.1-Flash is a 552-billion-parameter mixture-of-experts model with native multimodal visual understanding.
- The company retired V4-Flash and V4-Flash-Vision-Exp, while the older model names deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route requests to V4.1-Flash.
Selected from this article · 2026-09-10
Read on for the full pictureDeepSeek unveils V4.1-Flash as the smallest entry in a new architecture family

Chinese AI startup DeepSeek launched DeepSeek-V4.1-Flash on Thursday, September 10, 2026, calling it the smallest model in its new architecture family. According to a company statement, the model is designed for greater capability, faster inference, higher throughput and scaling to larger models. The release came as DeepSeek is starting to prepare for an initial public offering on Shanghai's tech-focused STAR Market.
Architecture, parameters and training footprint
V4.1-Flash is a 552-billion-parameter mixture-of-experts model with native multimodal visual understanding. DeepSeek said the model activates only 8 billion parameters during input and 16 billion during output, and supports both image and text inputs with an adjustable reasoning effort ranging from one to 100. The model uses a Causal Encoder-Decoder architecture with sparse attention and FP4 KV caching, bringing the global KV cache footprint down to 890 bytes per token and using one-quarter of the HBM and one-eighth of the SSD storage required by the previous generation.
The architecture consists of 20 encoder and 20 decoder layers, with a DeepSeek-ViT vision encoder that converts images into visual embeddings processed alongside text. An advanced MoE design pairs one shared expert with 384 routed experts, activating six routed experts per token. DeepSeek-V4.1-Flash was trained from scratch on a multimodal corpus containing 45 trillion tokens, with sparse attention trained at a 64K sequence length and the context window extended to 1 million tokens after training on 34 trillion tokens.
Context window, deprecations and price positioning
DeepSeek made V4.1-Flash available through its API with a context window of up to 1 million tokens. The company retired V4-Flash and V4-Flash-Vision-Exp, while the older model names deepseek-v4-flash and deepseek-v4-flash-vision-exp will temporarily route requests to V4.1-Flash. DeepSeek said multi-party testing found V4.1-Flash ahead of V4-Pro on performance, cost, speed and total completion time, prompting the phase-out of the older model. The company is pricing the new model aggressively, reportedly charging as little as a fraction of a cent per million tokens, intensifying the price battle between Chinese open-weight models and U.S. counterparts.
Competitive benchmarks and ecosystem adoption
DeepSeek claims V4.1-Flash outperforms Anthropic's Opus 5 on benchmarks including Terminal-Bench 3.0, DeepSWE v1.1, CyberGym and Automation-Best, and outperformed OpenAI's GPT-5.6 Sol on certain benchmarks. The model also surpassed Z.AI's GLM-5.3 in DeepSeek's testing; Z.AI had open-sourced a cheaper GLM-5.3-Flash in August at roughly one-tenth the price of GLM-5.3. Cambricon Technologies achieved "Day 0" support for V4.1-Flash, allowing the model to run immediately on its hardware upon release, and Tencent integrated the model through its WorkBuddy and CodeBuddy products alongside OpenCode as official partners.
Geopolitical backdrop and IPO preparation
The launch follows a U.S. government accusation that DeepSeek and five other Chinese AI companies used AI model distillation to extract capabilities from U.S. frontier models. Separately, Reuters reported that DeepSeek has hired CITIC Securities to prepare for a listing on Shanghai's STAR Market, with the process expected to begin this year, though the timing, size and valuation of the offering have not been finalized.
Follow-up signal
The next verifiable milestone is the formal start of DeepSeek's STAR Market listing process with CITIC Securities, expected to begin later in 2026, alongside any further pricing or benchmark disclosures for V4.1-Flash as the model is deployed by partners such as Cambricon Technologies and Tencent.
Share this article







