English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Alibaba Open-Sources Qwen3.8-2.4T-A95B: A Max-Class 2.4T MoE Flagship with Day-0 SiliconFlow Hosting and 9 Domestic AI Chips

Forum topic · 小凯 · 2026-08-15

Summary

On August 12, Alibaba's Qwen team released the full weights of Qwen3.8-2.4T-A95B, the open-weight counterpart of its Qwen3.8-Max commercial flagship. The model is a 2.4-trillion-parameter Mixture-of-Experts with roughly 95B parameters activated per token, 512 experts (11 activated per inference), and 92 hybrid attention layers (MLA plus sliding-window attention). Native context is 262,144 tokens, extendable via YaRN to about 1 million, with a recommended maximum output of 131,072 tokens. SiliconFlow launched the model Day-0 at $2.00/M input and $6.00/M output tokens, matching Grok 4.6 pricing. Nine domestic Chinese AI chips reportedly achieved Day-0 adaptation, including Ascend, Hygon, and Cambricon, shortening enterprise deployment paths. Benchmarks show strong reasoning (GPQA Diamond 92.6%, PaperBench 93.0%) but weaker long-horizon software engineering (DeepSWE v1.1 at 56.6, behind GLM-5.3's 66.9). The open weights are text-only with thinking mode forced on, and appear 10-15% behind the commercial Max API. This article summarizes the architecture, pricing, benchmark comparisons, and buying advice versus GLM-5.3, DeepSeek V4 Pro, and Grok 4.6.

On August 12, Alibaba's Qwen team officially open-sourced the weights of Qwen3.8-2.4T-A95B — 2.4 trillion total parameters, 95B activated per token, 512 experts with 11 activated per inference, and 92 hybrid attention layers. This is the first time Alibaba has fully released a Qwen-Max-class flagship: not a distilled, pruned, or trimmed version, but the open-weight equivalent of Qwen3.8-Max.

MoE Architecture: How 2.4T "Looks Like" Only 95B

The "A95B" in the name means Activation 95B — roughly 95 billion parameters are actually activated per token. Key specs:

  • Total parameters: 2.4 trillion
  • Activated parameters: ~95B per token
  • Experts: 512
  • Experts activated per inference: 11
  • Layers: 92 hybrid attention layers (MLA + sliding-window attention mix)
  • Native context: 262,144 tokens
  • Extended context: ~1,010,000 tokens via YaRN
  • Max output: 131,072 tokens
  • SiliconFlow launched the model Day-0, priced at $2.00/M input tokens and $6.00/M output tokens — nearly identical to Grok 4.6's $2/$6.

    There are differences from the commercial API version: the open-weight release is text-only (no vision or other multimodal input), and thinking mode is forced on (every response first generates a chain of thought, which cannot be disabled). For vision input or non-thinking mode, you need the Qwen3.8-Max commercial API.

    Day-0 Support on Domestic Compute: 9 Chinese AI Chips

    According to the GitHub AI Daily report, Qwen3.8-2.4T-A95B achieved Day-0 adaptation on 9 domestic AI chips on launch day. Major Chinese AI chips — Cambricon, Hygon, Ascend, Enflame, Iluvatar, Moore Threads, Biren, and others — can reportedly run this 2.4T model on day one without waiting for community ports.

    This may matter more than the parameter count. Previously, domestic open-source LLMs typically lagged 2–4 weeks before running on Ascend/Hygon/Cambricon. Day-0 adaptation means:

    1. Compute scheduling power shifts to domestic silicon. DeepSeek V4 Pro and GLM-5.3 previously ran on NVIDIA first, with domestic support added later; Qwen3.8 is the first to be "available on domestic compute from launch." 2. Shorter enterprise deployment paths. Government and state-owned customers can go from "wait for domestic adaptation" to "purchase domestic chips and deploy on release." 3. An earlier price war. SiliconFlow's Day-0 launch plus 9-chip adaptation means Qwen3.8 inference costs could quickly be pushed to $1/$3 or lower.

    Coding: Strong Reasoning, Weaker Long-Horizon Coding

    Qwen3.8 performs well on reasoning benchmarks: GPQA Diamond 92.6%, PaperBench 93.0%, Terminal-Bench 2.1 86.6%. But on DeepSWE v1.1 (long-horizon software engineering) it scores only 56.6, clearly behind GLM-5.3's 66.9 and DeepSeek V4 Pro's 62.7.

    Coding agent comparison:

  • Terminal Bench 2.1: Qwen3.8 86.6 / GLM-5.3 88.2 / DeepSeek V4 Pro 87.9 / Fable 5 88.0
  • DeepSWE v1.1: Qwen3.8 56.6 / GLM-5.3 66.9 / DeepSeek V4 Pro 62.7 / Fable 5 67.5
  • CyberGym (cybersecurity): Qwen3.8 78.1 / GLM-5.3 84.5 / Fable 5 83.8
Qwen3.8's strength is reasoning breadth (math, paper reproduction, OS operations); its weakness is end-to-end multi-file project coding. This is a notable gap versus the Max version's official positioning of "autonomous coding, deep research, end-to-end agent" — the open version may trail the commercial version by 10–15%.

How to Choose: Qwen3.8 vs GLM-5.3 vs DeepSeek V4 Pro

| Model | Input/Output ($/M tokens) | Notes | |---|---|---| | Qwen3.8-2.4T-A95B (SiliconFlow) | $2.00 / $6.00 | Open weights, 9-chip Day-0 | | GLM-5.3 (API) | TBD | Weights open-sourcing in two weeks | | DeepSeek V4 Pro (post 8/17 peak) | $1.27 / $3.80 | Still the cheapest flagship globally | | Grok 4.6 (re-billed >200K context) | $2.00 / $6.00 | Entire requests >200K tokens billed at $4/$12 |

One-line recommendation: choose Qwen3.8 for open-source self-hosting (available now) or wait for GLM-5.3 (two weeks); GLM-5.3 for the strongest open-source coding; DeepSeek V4 Pro for cheap-and-capable (lock in pricing before the 8/17 increase); Grok 4.6 for long-running agents + vision.

Three signals to watch over the next month: 1) whether domestic chip vendors publish Qwen3.8 inference benchmarks (especially Ascend 910C); 2) how many LoRA/QLoRA fine-tuned derivatives the community produces within 24 hours; 3) whether the capability gap between the Qwen3.8-Max commercial API and the open version continues to narrow (the distillation loss curve).

Sources: Qwen GitHub, ModelsLab Qwen3.8-2.4T docs, SiliconFlow announcement on X, ModelScope community Qwen3.8 page, GitHub AI Daily 8-15.

Tags

#qwen3.8#alibaba#open-source-llm#moe-architecture#siliconflow#domestic-ai-chips#benchmark#llm-pricing

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633490