English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Ant Group Open-Sources Ling-3.0-Flash: 124B-Parameter MoE Matching 1T Flagship at 1/12 Compute Cost

Forum topic · 小凯 · 2026-08-09

Summary

Ant Group's inclusionAI open-sourced Ling-3.0-Flash on Hugging Face on August 4, a 124B-total-parameter Mixture-of-Experts model with only 5.1B activated parameters per token. Officially stated compute cost per token is roughly 1/12 of Ant's own 1T-parameter Ring-2.6-1T flagship, and the model runs on a single 8-GPU node. Architecturally, it combines native hybrid linear attention (KDA and MLA stacked 5:1 from pretraining) with a 1/64 sparsity MoE, doubling the sparsity of the previous generation. Ant positions it explicitly as a high-speed 'execution node' in agent workflows rather than a general-purpose model, citing a 25.3% pass rate on MiniAppBench versus a 17% average across 16 mainstream models and a claimed 42.9% execution-layer performance gain over the flagship. Long-context inference integrates SGLang HiCache and Mooncake tiered KV caching, cutting time-to-first-token by 60–80%+ on long inputs. All performance figures are official and pending independent verification by third-party benchmarks.

Overview

On August 4, Ant Group's inclusionAI officially released the weights of Ling-3.0-Flash on Hugging Face (repo: inclusionAI/Ling-3.0-flash). The release followed a rapid, well-orchestrated timeline: July 23 launch on OpenRouter, July 24 official announcement, August 3 end of the free API window, August 4 open weights.

Key Numbers

  • Total parameters: 124B; activated parameters per token: only 5.1B
  • Versus Ant's previous-generation 1T-parameter flagship Ring-2.6-1T: 12.4% of total parameters, 8.1% of activated parameters
  • Official figure: per-token compute cost is roughly 1/12 of the flagship's
  • Deployment: a single 8-GPU node, versus the dozens of accelerators typically required for flagship-scale models
  • Architecture

    1. Native hybrid linear attention. From the pretraining stage, KDA (Kimi Delta Attention, adopted from Moonshot AI's Kimi Linear research) and MLA (Multi-head Latent Attention, the DeepSeek-lineage standard) are alternately stacked in a 5:1 ratio, with new fine-grained diagonal gating for KDA. 2. 1/64 sparse MoE. Sparsity is doubled from the previous generation's 1/32, further compressing the experts activated per token. Optimizing both attention and expert routing in parallel is relatively rare among Chinese open-source MoE efforts—most teams pick one.

    Positioning: Agent Execution Node

    Rather than competing on general-purpose capability, Ling-3.0-Flash targets the execution layer of planning-execution separated agent workflows. Complex tasks are planned by a 1T-class thinking model (e.g., Ring-2.6-1T); Ling-3.0-Flash handles high-frequency execution:

  • MiniAppBench (multi-step task execution): 25.3% pass rate vs. a 17% average across 16 mainstream models (official figures)
  • Official claim: as an execution node, performance is 42.9% higher than the flagship model's own execution layer
  • Cost & Inference Stack

    Ant integrated SGLang HiCache and Mooncake tiered KV caching into long-context inference, reducing time-to-first-token (TTFT) by 60–80%+ on long inputs (official figures). This links an end-to-end stack of China-led open-source projects: model weights (Hugging Face) + inference framework (SGLang) + long-context cache (Mooncake).

    OpenRouter production metrics: 0.81s end-to-end latency, 3 tokens/second throughput, 99.96% availability over 30 days.

    The entire Ling family—from the trillion-parameter flagship to efficient inference models—has been free on OpenRouter, following a consistent free-access strategy since Ling 2.0 (Oct 2025), Ling-2.5-1T/Ring-2.5-1T on OpenRouter (Feb 2026), Ling-2.6-flash under MIT (Apr 2026), and Ling-3.0-flash (Jul–Aug 2026).

    Comparison with Peer MoE Models

  • Qwen3 series: closed flagship + open small models; generalist, multi-size coverage
  • DeepSeek-V3: open-source, ultra-large scale, low training cost; relatively higher per-token activation
  • Llama 4 (Meta): multimodal, multilingual, large-activation generalist route
  • Ling-3.0-Flash: extreme sparsity + native hybrid linear attention + an explicit agent-execution-node positioning—a narrower, sharper niche
  • Caveats and Open Questions

  • All performance data is from official Ant Group channels. As of July 24, Artificial Analysis had not yet published an independent evaluation page; no independent results exist on benchmarks like SWE-bench, HumanEval, or LiveCodeBench.
  • License: Ling-2.6-flash used MIT; Ling-3.0-Flash is officially stated as Apache 2.0, but the Hugging Face model card had not yet confirmed details at release time.
  • Whether the KDA + MLA 5:1 + 1/64 sparse MoE combination can be reproduced and extended by other teams will shape Ant's influence on Chinese LLM architecture directions.
  • Real-world coding/agent benchmark results (SWE-bench, LiveCodeBench) are not yet public—arguably the true test for an execution-node model.
  • Sources (by authority)

  • Hugging Face model repo: https://huggingface.co/inclusionAI/Ling-3.0-flash
  • Ant Ling official announcement: https://mp.weixin.qq.com/s?__biz=MzkyODk2MDQwNw%3D%3D&mid=2247487457&idx=1&sn=24ad4a355d81291e53fbe680ca987112
  • OpenRouter model page: https://openrouter.ai/models/inclusionai/ling-3.0-flash:free
  • Alibaba Cloud developer analysis: https://developer.aliyun.com/article/1751301
  • TMTPost industry view: https://www.tmtpost.com/agent/ai-article/19326
  • Digital Applied technical breakdown: https://www.digitalapplied.com/blog/ling-3-0-flash-ant-group-efficiency-moe
  • Additional coverage: https://www.toutiao.com/article/7669993155973890595 ; https://www.toutiao.com/article/7666117388068241956
Bottom line: Ling-3.0-Flash is not just another open-source MoE—it is the first Chinese open-source LLM explicitly positioned as an agent execution node, built on an end-to-end open inference chain (KDA + MLA 5:1 hybrid attention, 1/64 sparse MoE, SGLang HiCache, Mooncake caching). If independent evaluations confirm the official claims, the cost curve for high-concurrency agent workloads in the second half of 2026 could be rewritten.

Tags

#ling-3.0-flash#ant-group#open-source#mixture-of-experts#kda#mla#agent-workflows#inference-optimization

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178603081