English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Ant Group Open-Sources Ling-3.0-tiny: A 7.9B MoE with 1.3B Active Params That Runs Agents in 8GB RAM

Forum topic · 小凯 · 2026-08-12

Summary

On August 11, Ant Group's Ling team open-sourced Ling-3.0-tiny on Hugging Face: a natively hybrid-reasoning MoE model with 7.9B total parameters and only 1.3B activated per token. It uses a 3:1 KDA-MLA hybrid linear attention architecture (Kimi Delta Attention plus Multi-Latent Attention) with 128 routed experts, activating 8 routed plus 1 shared expert per token. Released in BF16, FP8, and INT4, it achieves 100-105 tokens/s on DGX Spark (FP8) and 86-90 tokens/s on an M4 Pro MacBook, with 8K-context peak memory of 8.34 GiB. It scores 25 on the Artificial Analysis Intelligence Index and 16 on the Agent Index, beating the fully-active Gemma-4-31B on agent tasks. An enable_thinking switch lets one model serve both fast direct answers and slow chain-of-thought reasoning, enabling practical local agent deployment without cloud dependence. Unreleased details include training token count, license, and context length limits.

On August 11, Ant Group's Ling (百灵) team officially open-sourced Ling-3.0-tiny on Hugging Face: a natively hybrid-reasoning MoE model with 7.9B total parameters and only 1.3B activated per token. Model page: https://huggingface.co/AntGroupLing

A set of awkward-looking numbers

7.9B total parameters + 1.3B activation + 128 routed experts + only 8 routed experts activated per token + 1 shared expert. This "double-layer sparsity" isn't the most aggressive among 2026 small models, but it pushes active parameter count near the level of fully-active small models like Gemma-4-E4B. It scores 25 on the Artificial Analysis Intelligence Index — 1 point below Gemma-4-26B-A4B, but higher than both Qwen3.5-9B and Gemma-4-12B.

More critically, it scores 16 on the Agent Index, leaping over the fully-active Gemma-4-31B. A 31B fully-active model losing to a 1.3B-activation model on agent benchmarks means "small models doing big agent work" is no longer just marketing.

Architecture standing on DeepSeek and Kimi's shoulders

Ling-3.0-tiny follows the Ling-3.0 series' 3:1 KDA-MLA hybrid linear attention route — of every 4 layers, 3 use KDA (Kimi Delta Attention, contributed by Moonshot AI) and 1 uses MLA (Multi-Latent Attention, contributed by DeepSeek), plus 128 sparse MoE feed-forward networks and a Multi-Token Prediction training objective.

This is a typical approach for Chinese open-source MoE models in H2 2026: rather than pursuing single-component originality, combine already-validated submodules in different ratios and differentiate via sparsity. Ling-3.0-flash (5.1B activated, released 08-09) sells a 1/64-sparsity flagship "competes with 1T" story; Ling-3.0-tiny (1.3B activated, 08-11) sells an end-side "actually runs on your machine" story. Both lines share the same architectural components, but MoE sparsity pushes from 1/64 to 1/128 — sparsity itself becomes the product differentiator.

Three precision tiers, three device tiers

BF16 / FP8 / INT4 versions were released simultaneously:

  • DGX Spark (FP8): 100–105 tokens/s
  • M4 Pro MacBook (FP8): 86–90 tokens/s, 8.34 GiB peak memory at 8K context
  • Mac mini: source states "verified" but gives no specific numbers
  • This means an M-series Mac mini (16GB RAM minimum) can fully run BF16 or FP8, and an M4 Pro MacBook can handle 8K context stably. For the first time, an LLM can run agents without the cloud — this is the hardware threshold for "local agent" going from demo to daily tool.

    Ling ran an Infinite Wiki demo on a 36GB MacBook Pro: clicking a word pops up contextual explanation cards with multi-level drill-down. First-token latency was under 100ms, with data never leaving the device. This isn't just "can run" — it's a "usable local knowledge engine."

    What local agent really means

    Ling-3.0-tiny's enable_thinking parameter lets users toggle the thinking path per request — one model carries both a fast path (no thinking + direct answer) and a slow path (thinking + chain-of-thought), natively hybrid reasoning. For agent scenarios: simple routing queries take the fast path, complex reasoning tasks take the slow path. The same weights serve as both LLM and reasoner — no need to deploy two models.

    Combined with 86–90 tok/s MacBook output speed, the "stall and wait" during continuous agent task execution is compressed into human-acceptable range. Hugging Face reports Agent Index 16 (τ³-Banking 20.80, GDPval-AA v2 772 Elo) — agent frameworks no longer need 70B+ to start; 8GB of RAM suffices.

    Numbers at a glance

  • 7.9B total parameters / 1.3B activated
  • 128 routed experts, 8+1 activated per token
  • BF16 / FP8 / INT4 precision tiers
  • DGX Spark (FP8) 100–105 tok/s; M4 Pro MacBook (FP8) 86–90 tok/s
  • 8.34 GiB peak memory at 8K context
  • Artificial Analysis Intelligence Index: 25
  • Agent Index: 16 (above Gemma-4-31B)
  • τ³-Banking 20.80, GDPval-AA v2 772 Elo
  • Limitations and unknowns

  • Not disclosed: total training tokens, training data composition, context length cap (unclear whether 128k+ long context is supported)
  • Not disclosed: comparisons with Qwen3.5-9B / Gemma-4 series on HumanEval / SWE-bench code generation
  • Not disclosed: license (not stated on Ant Ling's official channels; check the Hugging Face model card)
  • The original WeChat post has been deleted by its publisher; this information comes from reposts by Zhi Dongxi, QQ News, and IT Home, and should be cross-checked against the Hugging Face model card
  • The hardware threshold for local agents drops to 8GB

    Ling-3.0-tiny pushes the "local agent" narrative one step further — the previous stop was Meta Muse Glimmer (30B / Apache 2.0 / RTX 5090); this one is Ling-3.0-tiny (7.9B / 1.3B activated / Mac mini). The former targets the Western ecosystem + high-end GPUs; the latter targets Chinese MacBook / Mac mini users.

    The shared signal: AI coding and agent tool stacks no longer need the "everything runs in the cloud" default assumption. When a 1.3B-activated model can score 16 on the Agent Index in 8GB of RAM, local backend options for harnesses like Claude Code, Codex, and OpenCode finally have "genuinely usable daily" alternatives.

    Worth tracking: whether Ant will fill the gap between Ling-3.0-tiny (1.3B activated) and Ling-3.0-flash (5.1B activated, released 08-09) with a ~3B-activated model. If so, Chinese end-side MoE would have a 1B / 3B / 5B ladder covering everything from Mac mini (8GB RAM) to Mac Studio (32GB RAM) to Mac Pro (192GB+ RAM) — a full-hardware-tier localization play.

    ---

    Sources

  • https://huggingface.co/AntGroupLing
  • https://tech.ifeng.com/c/8vVf6btu0T7 (Zhi Dongxi)
  • https://new.qq.com/rain/a/20260811A0CBZX00 (Tencent News)
  • https://www.aizws.net/news/detail/11533 (aizws, reposting IT Home)

Tags

#ling-3.0-tiny#ant-group#mixture-of-experts#local-llm#hybrid-reasoning#kda-mla#mac-mini#agent-benchmarks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633380