What changed

Person's hand inserting a USB flash drive into a laptop. Technology and data transfer concept.

Alibaba released the text-only Qwen3.8-2.4T-A95B open-weight checkpoint behind its managed Qwen3.8-Max service on Hugging Face, the first time a Max-class model in the Qwen family has been made publicly downloadable. The Hangzhou-based company had previously reserved open-weight licences for its smaller models and kept the Max range proprietary, so the release marks a shift from its earlier open-weights policy.

The model and how it is built

The checkpoint carries 2.4 trillion total parameters but activates only about 95 billion per operation through a sparse mixture-of-experts architecture paired with a hybrid attention mechanism. Each of the 23 repeating blocks contains three Gated DeltaNet-plus-MoE units followed by one Gated Attention-plus-MoE unit, mixing linear and standard gated attention. The model card lists a native 262,144-token context length, extendable to 1,010,000 tokens, and exposes reasoning effort and preserve-thinking controls. NVIDIA has stated that multi-node, data-centre-scale accelerated systems are required, reporting throughput of more than 4,000 tokens per second per GPU and more than 350 tokens per second per user on a GB300 NVL72 system at FP8 precision.

Benchmarks and competitive positioning

Independent leaderboards place Qwen3.8-Max second on Vision Arena and fifth on Text Arena, trailing only Anthropic's Claude Fable 5 in overall performance. At 2.4 trillion parameters it approaches Moonshot AI's Kimi K3 at 2.8 trillion parameters, which is the largest available open-weight model and likewise imposes paid conditions for large-scale commercial use. Both Chinese systems are more than twice the size of an open-weight model with at least 1 trillion parameters that Nvidia is reportedly developing, while Meta said it would release the weights of its Muse Spark 1.2 model in the coming weeks.

Licensing and access terms

The checkpoint is governed by the Qwen3.8-Max License rather than Apache 2.0 and requires prominent model-name display for commercial products or services above 100 million monthly active users or $20 million in monthly revenue. A separate Qwen license is required once aggregate revenue exceeds $50 million during any consecutive 12 months, and the open weights may be used freely only for internal purposes as long as the software, outputs and underlying capabilities are not made available to third parties. Alongside the download, Alibaba is offering paid API access at 12 yuan per million input tokens and 36 yuan per million output tokens domestically, and $2 per million input tokens and $6 per million output tokens internationally.

Open versus managed service

The downloadable open-weight checkpoint is text-only and runs only in thinking mode, whereas the managed Qwen3.8-Max service adds vision input, a non-thinking mode, and built-in tools that the open version excludes. Storage, loading, expert placement, interconnect bandwidth and quantization remain significant constraints even though only about 4% of parameters activate per token. The deployment footprint has been characterized by analysts as a credible way for Alibaba to undercut US frontier pricing rather than merely match capability.

What remains uncertain

It is not yet confirmed how organisations will reconcile the open checkpoint's licensing thresholds with hosted-service deployments, and benchmark positioning against Anthropic's Claude Fable 5 and Moonshot's Kimi K3 has not yet been independently audited beyond leaderboard placements.

Share this article

FacebookX

2 sources

Sources