Catch up on the essentials
- Alibaba's Qwen team released Qwen3.8-Omni-Flash, a native omnimodal model that accepts text, images, audio and video in a single workflow and supports a one-million-token context window.
- Qwen framed the headline capability as agentic perception, in which the model decides which parts of a recording to examine before gathering evidence across coarse-to-fine passes rather than sampling every frame at fixed intervals.
- Measured per hour of media processed, Qwen said audio-input costs fall by about 98% and combined audio-video costs by roughly 89% compared with Qwen3.5-Omni-Plus, a reduction it attributes partly to the lower sampling rate that agentic perception allows.
Selected from this article · 2026-09-20
Read on for the full pictureQwen3.8-Omni-Flash architecture and long-context design

Alibaba's Qwen team released Qwen3.8-Omni-Flash, a native omnimodal model that accepts text, images, audio and video in a single workflow and supports a one-million-token context window. According to Qwen, the model is now available through the Qwen AI platform, and built directly on the Qwen3.8-Flash-Next architecture whose open weights shipped in August 2026, rather than assembling a text model with separate perception encoders. Internally, the system uses a sparse mixture-of-experts design with roughly 125 billion total parameters, of which about 6 billion activate per token across 512 experts, a configuration Qwen says is intended to hold down inference costs while keeping a large pool of specialised parameters available. The reported maximum input sits close to one million tokens, around 991,000, and the model accepts video files up to two hours and audio up to three hours in a single request.
Agentic perception and benchmark performance
Qwen framed the headline capability as agentic perception, in which the model decides which parts of a recording to examine before gathering evidence across coarse-to-fine passes rather than sampling every frame at fixed intervals. On an internal long-video benchmark, Qwen reported that this approach lifted accuracy from 63.4 to 67.8 while reducing the tokens needed per query by about 46%, from roughly 145,700 to 79,100. The company said Qwen3.8-Omni-Flash improved its average score by more than 25% across 29 evaluations, and by more than 26% across roughly 30 evaluations, compared with Qwen3.5-Omni-Plus, with the largest gains on agent-oriented tests combining perception, planning and tool use. Qwen also said the model's audio-visual performance approaches Google's Gemini 3.8 Flash and that its overall audio performance exceeds it, while acknowledging Gemini still leads on some pure video-reasoning benchmarks. Independent testing of these claims was not available in the source set.
Pricing, regional availability and open-source tooling
API pricing for Qwen3.8-Omni-Flash is set at $0.15 per million input tokens and $0.47 per million output tokens, with cached input tokens billed at $0.016 per million, and the API input price has been reduced to as low as RMB 0.8 per million tokens. Measured per hour of media processed, Qwen said audio-input costs fall by about 98% and combined audio-video costs by roughly 89% compared with Qwen3.5-Omni-Plus, a reduction it attributes partly to the lower sampling rate that agentic perception allows. The model is live through Qwen Chat, Alibaba Cloud Model Studio and the QwenCloud API across six regions, including Singapore, Tokyo and Frankfurt, though only as a hosted service with no open weights released for this version. Alongside the model, Alibaba open-sourced a companion toolkit called Qwen-MM-Plugins under an Apache 2.0 license, designed to give coding-agent harnesses, including Claude Code, Codex, Gemini CLI and Qwen Code, native support for reading images, video frames and audio locally rather than relying entirely on the hosted API, and it introduced Qwen-Live Harness for long-running and real-time multimodal workflows.
Targeted workflows from meetings to film production
Qwen positioned the model for a set of end-to-end media and productivity workflows. For long videos, the model can analyze characters, camera shots, lighting and sound based on user prompts, and Qwen said the approach reduced token use by around 45.7% on OmniVideoBench. For meetings, Qwen3.8-Omni-Flash supports up to one hour of audio-visual input, identifying speakers, transcribing discussions, generating minutes, extracting action items, analyzing project risks, and, when connected to tools, sending emails, organizing tasks or beginning coding based on meeting requirements. For media production, Qwen said the model can analyze music before planning music videos and run a video translation workflow covering transcription, translation, voice cloning, dubbing, audio mixing and final review, while a separate agent handles plot extraction, script planning, voiceover, music, editing, rendering and quality checks for full-length films.
Realtime variant and language coverage
Qwen also introduced Qwen3.8-Omni-Flash-Realtime, which processes live audio and video while responding and using tools, and the company said it can combine visual information with spatial sound to determine where sounds are coming from and assist with localization and navigation. The realtime variant supports speech recognition in 74 languages, including Urdu and Punjabi, with speech generation covering 29 languages. Qwen has not yet set a release date for the real-time streaming counterpart, which it previewed as a live harness for continuous voice and video interaction.
Share this article







