On August 12, Alibaba's Qwen team officially open-sourced the weights of Qwen3.8-2.4T-A95B — 2.4 trillion total parameters, 95B activated per token, 512 experts with 11 activated per inference, and 92 hybrid attention layers. This is the first time Alibaba has fully released a Qwen-Max-class flagship: not a distilled, pruned, or trimmed version, but the open-weight equivalent of Qwen3.8-Max.
MoE Architecture: How 2.4T "Looks Like" Only 95B
The "A95B" in the name means Activation 95B — roughly 95 billion parameters are actually activated per token. Key specs:
- Total parameters: 2.4 trillion
- Activated parameters: ~95B per token
- Experts: 512
- Experts activated per inference: 11
- Layers: 92 hybrid attention layers (MLA + sliding-window attention mix)
- Native context: 262,144 tokens
- Extended context: ~1,010,000 tokens via YaRN
- Max output: 131,072 tokens
- Terminal Bench 2.1: Qwen3.8 86.6 / GLM-5.3 88.2 / DeepSeek V4 Pro 87.9 / Fable 5 88.0
- DeepSWE v1.1: Qwen3.8 56.6 / GLM-5.3 66.9 / DeepSeek V4 Pro 62.7 / Fable 5 67.5
- CyberGym (cybersecurity): Qwen3.8 78.1 / GLM-5.3 84.5 / Fable 5 83.8
SiliconFlow launched the model Day-0, priced at $2.00/M input tokens and $6.00/M output tokens — nearly identical to Grok 4.6's $2/$6.
There are differences from the commercial API version: the open-weight release is text-only (no vision or other multimodal input), and thinking mode is forced on (every response first generates a chain of thought, which cannot be disabled). For vision input or non-thinking mode, you need the Qwen3.8-Max commercial API.
Day-0 Support on Domestic Compute: 9 Chinese AI Chips
According to the GitHub AI Daily report, Qwen3.8-2.4T-A95B achieved Day-0 adaptation on 9 domestic AI chips on launch day. Major Chinese AI chips — Cambricon, Hygon, Ascend, Enflame, Iluvatar, Moore Threads, Biren, and others — can reportedly run this 2.4T model on day one without waiting for community ports.
This may matter more than the parameter count. Previously, domestic open-source LLMs typically lagged 2–4 weeks before running on Ascend/Hygon/Cambricon. Day-0 adaptation means:
1. Compute scheduling power shifts to domestic silicon. DeepSeek V4 Pro and GLM-5.3 previously ran on NVIDIA first, with domestic support added later; Qwen3.8 is the first to be "available on domestic compute from launch." 2. Shorter enterprise deployment paths. Government and state-owned customers can go from "wait for domestic adaptation" to "purchase domestic chips and deploy on release." 3. An earlier price war. SiliconFlow's Day-0 launch plus 9-chip adaptation means Qwen3.8 inference costs could quickly be pushed to $1/$3 or lower.
Coding: Strong Reasoning, Weaker Long-Horizon Coding
Qwen3.8 performs well on reasoning benchmarks: GPQA Diamond 92.6%, PaperBench 93.0%, Terminal-Bench 2.1 86.6%. But on DeepSWE v1.1 (long-horizon software engineering) it scores only 56.6, clearly behind GLM-5.3's 66.9 and DeepSeek V4 Pro's 62.7.
Coding agent comparison:
How to Choose: Qwen3.8 vs GLM-5.3 vs DeepSeek V4 Pro
| Model | Input/Output ($/M tokens) | Notes | |---|---|---| | Qwen3.8-2.4T-A95B (SiliconFlow) | $2.00 / $6.00 | Open weights, 9-chip Day-0 | | GLM-5.3 (API) | TBD | Weights open-sourcing in two weeks | | DeepSeek V4 Pro (post 8/17 peak) | $1.27 / $3.80 | Still the cheapest flagship globally | | Grok 4.6 (re-billed >200K context) | $2.00 / $6.00 | Entire requests >200K tokens billed at $4/$12 |
One-line recommendation: choose Qwen3.8 for open-source self-hosting (available now) or wait for GLM-5.3 (two weeks); GLM-5.3 for the strongest open-source coding; DeepSeek V4 Pro for cheap-and-capable (lock in pricing before the 8/17 increase); Grok 4.6 for long-running agents + vision.
Three signals to watch over the next month: 1) whether domestic chip vendors publish Qwen3.8 inference benchmarks (especially Ascend 910C); 2) how many LoRA/QLoRA fine-tuned derivatives the community produces within 24 hours; 3) whether the capability gap between the Qwen3.8-Max commercial API and the open version continues to narrow (the distillation loss curve).
Sources: Qwen GitHub, ModelsLab Qwen3.8-2.4T docs, SiliconFlow announcement on X, ModelScope community Qwen3.8 page, GitHub AI Daily 8-15.