Overview
On August 4, Ant Group's inclusionAI officially released the weights of Ling-3.0-Flash on Hugging Face (repo: inclusionAI/Ling-3.0-flash). The release followed a rapid, well-orchestrated timeline: July 23 launch on OpenRouter, July 24 official announcement, August 3 end of the free API window, August 4 open weights.
Key Numbers
- Total parameters: 124B; activated parameters per token: only 5.1B
- Versus Ant's previous-generation 1T-parameter flagship Ring-2.6-1T: 12.4% of total parameters, 8.1% of activated parameters
- Official figure: per-token compute cost is roughly 1/12 of the flagship's
- Deployment: a single 8-GPU node, versus the dozens of accelerators typically required for flagship-scale models
- MiniAppBench (multi-step task execution): 25.3% pass rate vs. a 17% average across 16 mainstream models (official figures)
- Official claim: as an execution node, performance is 42.9% higher than the flagship model's own execution layer
- Qwen3 series: closed flagship + open small models; generalist, multi-size coverage
- DeepSeek-V3: open-source, ultra-large scale, low training cost; relatively higher per-token activation
- Llama 4 (Meta): multimodal, multilingual, large-activation generalist route
- Ling-3.0-Flash: extreme sparsity + native hybrid linear attention + an explicit agent-execution-node positioning—a narrower, sharper niche
- All performance data is from official Ant Group channels. As of July 24, Artificial Analysis had not yet published an independent evaluation page; no independent results exist on benchmarks like SWE-bench, HumanEval, or LiveCodeBench.
- License: Ling-2.6-flash used MIT; Ling-3.0-Flash is officially stated as Apache 2.0, but the Hugging Face model card had not yet confirmed details at release time.
- Whether the KDA + MLA 5:1 + 1/64 sparse MoE combination can be reproduced and extended by other teams will shape Ant's influence on Chinese LLM architecture directions.
- Real-world coding/agent benchmark results (SWE-bench, LiveCodeBench) are not yet public—arguably the true test for an execution-node model.
- Hugging Face model repo: https://huggingface.co/inclusionAI/Ling-3.0-flash
- Ant Ling official announcement: https://mp.weixin.qq.com/s?__biz=MzkyODk2MDQwNw%3D%3D&mid=2247487457&idx=1&sn=24ad4a355d81291e53fbe680ca987112
- OpenRouter model page: https://openrouter.ai/models/inclusionai/ling-3.0-flash:free
- Alibaba Cloud developer analysis: https://developer.aliyun.com/article/1751301
- TMTPost industry view: https://www.tmtpost.com/agent/ai-article/19326
- Digital Applied technical breakdown: https://www.digitalapplied.com/blog/ling-3-0-flash-ant-group-efficiency-moe
- Additional coverage: https://www.toutiao.com/article/7669993155973890595 ; https://www.toutiao.com/article/7666117388068241956
Architecture
1. Native hybrid linear attention. From the pretraining stage, KDA (Kimi Delta Attention, adopted from Moonshot AI's Kimi Linear research) and MLA (Multi-head Latent Attention, the DeepSeek-lineage standard) are alternately stacked in a 5:1 ratio, with new fine-grained diagonal gating for KDA. 2. 1/64 sparse MoE. Sparsity is doubled from the previous generation's 1/32, further compressing the experts activated per token. Optimizing both attention and expert routing in parallel is relatively rare among Chinese open-source MoE efforts—most teams pick one.
Positioning: Agent Execution Node
Rather than competing on general-purpose capability, Ling-3.0-Flash targets the execution layer of planning-execution separated agent workflows. Complex tasks are planned by a 1T-class thinking model (e.g., Ring-2.6-1T); Ling-3.0-Flash handles high-frequency execution:
Cost & Inference Stack
Ant integrated SGLang HiCache and Mooncake tiered KV caching into long-context inference, reducing time-to-first-token (TTFT) by 60–80%+ on long inputs (official figures). This links an end-to-end stack of China-led open-source projects: model weights (Hugging Face) + inference framework (SGLang) + long-context cache (Mooncake).
OpenRouter production metrics: 0.81s end-to-end latency, 3 tokens/second throughput, 99.96% availability over 30 days.
The entire Ling family—from the trillion-parameter flagship to efficient inference models—has been free on OpenRouter, following a consistent free-access strategy since Ling 2.0 (Oct 2025), Ling-2.5-1T/Ring-2.5-1T on OpenRouter (Feb 2026), Ling-2.6-flash under MIT (Apr 2026), and Ling-3.0-flash (Jul–Aug 2026).