On the night of August 12 (Beijing time), Alibaba Cloud's ModelScope community announced that the Qwen team has fully released the weights of Qwen3.8-2.4T-A95B — the first Qwen-Max-class model to be completely open-sourced. Key specs: 2.4 trillion total parameters, 95B activated per token, native 262,144-token (256K) context expandable to 1,010,000 tokens (1M), text-only with forced thinking mode. Weights were simultaneously published to four inference stacks: Hugging Face Transformers, vLLM, SGLang, and TokenSpeed.
How It Compares: The Scale of This Release
| Model | Total Params | Active Params | Native Context | Open / Closed | |-------|-------------|---------------|----------------|---------------| | Qwen3.8-2.4T-A95B | 2.4T | 95B | 256K (ext. 1M) | Open | | Moonshot Kimi K3 | 2.8T | 104B | 1M+ | Semi-open (mainly distilled) | | DeepSeek V4-Pro | 1.6T | ~49B | 1M | Semi-open (partial architecture) | | Meta Llama 4 Behemoth | ~2T | ~100B | 128K | Closed | | xAI Grok 4.6 | 1.5T | 8–16 | 256K | Closed | | Claude Fable 5 Max | Undisclosed | — | 200K | Closed | | Microsoft MAI-Thinking-1 | 1T | 35B | 256K | Closed |
The takeaway: by August 2026, the scale gap between open and closed model camps has shrunk from 10× to roughly 1.5×. Open-source models are undergoing 'de-budgetization' — they used to default to 'mini versions of closed flagships'; this time the Max-class model itself was opened.
Three Architecture Details Worth Noting
1. 512 experts, 10 routed + 1 shared per token — a new MoE tier for 2026. Sparse at 1/218, among the densest expert pools yet sparsest activation in Chinese open MoEs. More expressible expert combinations with lower per-inference compute, but routing is a combinatorial challenge (~5.2×10¹¹ combinations).
2. 92 layers: 69 linear attention + 23 full attention — continuing Qwen3.5's hybrid attention approach. Linear attention compresses KV-cache memory while full-attention layers preserve precision, making 256K/1M contexts feasible within limited VRAM.
3. Multi-step Multi-Token Prediction (MTP) — building on the DeepSeek V3/R1 approach, with a built-in MTP draft head usable directly by vLLM, SGLang, and TokenSpeed for speculative decoding.
The Awkward Numbers: Who Can Actually Run It
| Format | Size | Hardware | |--------|------|----------| | BF16 full precision | ~4.9TB | DGX-class (H100/RTX PRO), 8 GPUs minimum | | FP8 quantized | ~2.5TB | H100/H200, 1–8 GPUs | | Unsloth TQ1_0 (1-bit) | <400GB | Single RTX 4090 / Mac Studio (M2 Ultra 192GB) | | 4-bit GPTQ | ~1.2TB | Single RTX 5090 / Mac Pro |
Qwen officially recommends FP8. The 1-bit variant is an extreme compression play with low throughput and degraded long-context quality. After this release, the narrative shifts: weights are open, but the hardware bar returns to the data center — you can study the code, do research, and fine-tune quantized copies, but not run the full model. This is a 'sovereign AI' path, not an 'independent developer' path.
Forced Thinking Mode: The Open Release's 'Capability Seal'
The open model requires thinking mode — all responses wrapped in thinking tags, no non-thinking option, no vision input. The cloud Qwen3.8-Max adds vision, a non-thinking fast mode, built-in tools (search, calculator, code execution), and default 1M context. The open release is a 'semi-finished core': researchers get language + long-context + agent foundations; product builders pay for the cloud version. Qwen wins both ways — community from open weights, revenue from the cloud tier.
A Piece of the H2 2026 'Sovereign AI' Puzzle
- 8/08: NVIDIA Cosmos 3 foundation model (1.3B data points, OpenMDW 1.1 commercial license)
- 8/08: Unitree Robotics IPO on the STAR Market at 219× PE
- 8/09: Ant Group Ling-3.0-flash open-sourced (5.1B active / 124B total, SGLang HiCache + Mooncake)
- 8/11: Ant Ling-3.0-tiny open-sourced (1.3B active / 7.9B total, 128 routed experts)
- 8/12: Alibaba Qwen3.8-2.4T-A95B open-sourced
- 8/12: NVIDIA Nemotron 4 trillion-parameter R&D info leak
- No vision input, no native tool calls — text-only with forced thinking
- PaperBench / IFBench scores are self-reported; no independent third-party verification yet
- 512-expert routing learning curve unverified; fine-tuning reproduction costs are non-trivial
- Commercial license not clearly stated — only a pointer to a Qwen3.8-specific license file
- Hardware barrier excludes individual developers except via extreme quantization
- Roughly tier-matched with Kimi K3 on total params; Qwen's edge is 'fully open' plus 'native 256K + forced thinking' positioning
- https://modelscope.cn/Qwen
- https://finance.sina.cn/tech/2026-08-13/detail-ininarri3875962.d.html
- https://www.163.com/dy/article/L46N0SML0556I485.html
- https://www.donews.com/news/detail/8/6668953.html
- https://setupai.cc/update/alibaba-opens-1ea180b0
- https://wpnews.pro/news/alibaba-releases-open-qwen3-8-with-2-4t-total-95b-active-parameters
Previously Qwen's largest open model was Qwen3-235B (235B total / 22B active), well below the closed Qwen3.5-Plus flagship at 720B. Opening 2.4T locks the open-vs-closed ceiling at Max parity. Implications: DeepSeek must decide whether V5 goes fully open; Moonshot's Kimi K3 lags on full openness (mostly distilled); Meta likely stays dual-track; closed vendors must ask whether 'closed = premium' still holds.
Limitations and Unknowns
What First-Time Max-Class Open-Sourcing Means
Individually, each number extends an existing trend; combined, they mark an inflection point: Qwen decided to fully open a Max-class model. Competition in H2 2026 shifts from 'whose activated parameter count is smaller and cheaper to run' to 'whose total scale is bolder to open.' DeepSeek, Kimi, MiniMax, ByteDance, and Zhipu all now face the same pressure test: will you open your Max-class models too? This was unthinkable in 2025. The move is less a technical breakthrough than a commercial posture — the open-source camp is now pressuring closed-source players.
---
Sources