English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MiniCPM5-2B: ModelBest's On-Device Agent Base Model Tops Sub-4B Rankings with Day-0 Support on 9 Chips

Forum topic · 小凯 · 2026-07-20

Summary

On July 19, 2026, at WAIC 2026, ModelBest (Mianbi Intelligence) and OpenBMB launched MiniCPM5-2B, a 2B-parameter on-device model codenamed 'Little Cannon'. It ranks first globally among sub-4B models on the AA-Index leaderboard with an average score of 54.26, beating even Liquid's 8B sparse LFM2.5-8B-A1B (49.15) as well as Qwen3.5-2B, Qwen3-1.7B, and Gemma-4-E2B-it. Key agent benchmarks include 90.35 on τ²-Bench Telecom, 60.95 on AIME-2026, and 69.60 on ZebraLogic. The model features hybrid thinking, 512K context length, tool calling, deep search, and code generation. It achieved Day-0 adaptation across nine chip platforms, including Huawei Ascend, Hygon, Kunlunxin, MetaX, Moore Threads, T-Head, Tsingmicro, Iluvatar CoreX, and NVIDIA, via the FlagOS unified software stack. ModelBest also announced a strategic MOU with Intel, demonstrating a 35B model running locally on an Intel 18A Core Ultra platform at 180 TOPS. An UltraX data-governance pipeline uses a 0.6B model for row-level data edits. Weights are expected to be open-sourced on ModelScope and Hugging Face.

MiniCPM5-2B: Sub-4B Global No.1 + Day-0 Support on 9 Chips — ModelBest Pushes the On-Device Agent Base Model Price to the Floor

Source: ModelBest / Jiazi Guangnian URL: https://mp.weixin.qq.com/s/rjFxrUylyGMqa5QtgypCdw Date: 2026-07-19 (WAIC 2026)

The Announcement

On July 19, ModelBest (Mianbi Intelligence), together with OpenBMB, released the on-device model MiniCPM5-2B (codename "Little Cannon") at WAIC:

  • AA-Index leaderboard: ranked global No.1 among sub-4B models with 17 points
  • Average score 54.26: first among sub-4B base models in overall capability
  • Day-0 adaptation on 9 chips: Huawei Ascend, Hygon, Kunlunxin, MetaX, Moore Threads, T-Head (Alibaba), Tsingmicro, Iluvatar CoreX, NVIDIA
  • Hybrid thinking + 512K context plus tool calling, deep search, and code generation
  • To be open-sourced: model weights plus per-chip adaptation images will be released on ModelScope and Hugging Face
  • Benchmark Analysis

    Beats All Same-Size Rivals

    | Model | Params | AA-Index Avg | |---|---|---| | MiniCPM5-2B | 2B | 54.26 | | LFM2.5-8B-A1B (8B sparse) | ~1B active | 49.15 | | Qwen3.5-2B | 2B | 41.66 | | Qwen3-1.7B | 1.7B | 40.89 | | Gemma-4-E2B-it | 2B | 37.62 |

    Note: LFM2.5-8B-A1B is Liquid's 8B sparse architecture (~1B active). ModelBest's dense 2B beats an 8B sparse model — this is a measured leaderboard result, not a slogan about parameter efficiency.

    Agent Capability Is the Real Focus

    | Task | MiniCPM5-2B | Notes | |---|---|---| | AA-LCR long-context understanding | 42.33 | 16 points ahead of second place | | τ²-Bench Telecom Agent | 90.35 | 4.97 points ahead of LFM2.5-8B-A1B | | AIME-2026 math reasoning | 60.95 | Qwen3-1.7B: 36.88 / Gemma-4-E2B-it: 32.71 | | ZebraLogic logical reasoning | 69.60 | First among compared models |

    τ²-Bench Telecom evaluates multi-step customer-service agent workflows. A score of 90.35 means a 2B on-device model can independently handle telecom-operator-level multi-step service flows — a direct enabler for real process-oriented tasks (not just Q&A) on phones and in-car assistants.

    Day-0 on 9 Chips: Cloud-to-Edge Coverage Across Architectures

    Via the FlagOS community's unified software stack (vLLM-plugin-FL + SGLang-plugin-FL), ModelBest completed synchronized adaptation on 9 AI chips:

    | Category | Chips | |---|---| | Domestic Chinese AI compute | Huawei Ascend, Hygon, Kunlunxin, MetaX, Moore Threads, T-Head, Tsingmicro, Iluvatar CoreX | | International GPU | NVIDIA | | Edge platform | ARM |

    ModelBest explicitly states: after FlagOS optimization, some domestic chips match or exceed the NVIDIA baseline in per-unit-compute inference performance — with AA-Index benchmarks run under multi-platform alignment.

    Intel 18A Strategic Partnership: 35B on 180 TOPS

    ModelBest also signed an MOU with Intel for joint optimization of on-device LLMs and chip platforms. Baseline figure: Intel's 18A process Core Ultra platform, at 180 TOPS, supports local deployment of a 35B-parameter model — a milestone number for 2026 on-device NPUs. Adaptations for AMD, MediaTek, and Qualcomm edge platforms are also in progress.

    UltraX Data Governance: Row-Level Edits by a 0.6B Model

    As the internet data dividend slows, ModelBest proposes tiered data governance (UltraX): a 0.6B model performs row-level additions, deletions, and edits on L1 web data, upgrading data governance from "traditional cleaning" to "executable editing programs" — the first time a model lab has engineered governance at this granularity.

    Full Agent Training Pipeline

  • Agent Midtraining at 200B-token scale
  • Agent SFT on millions of high-quality trajectories
  • Scalable Agent RL alignment
  • The 200B here refers to midtraining specialized for agent capability, mapping directly to the positioning of the model as an on-device agent base for local AI assistants on phones, PCs, and smart cockpits.

    Why It Matters

    1. AA-Index No.1 is verifiable — a third-party public leaderboard; the 54.26 average beats the 8B sparse LFM2.5. 2. τ²-Bench Telecom 90.35 is a key on-device agent milestone — a 2B model completing operator-grade multi-step flows means phone assistants can do more than Q&A. 3. Day-0 on 9 chips — against the backdrop of US export controls and domestic substitution, this multi-chip ecosystem reduces NVIDIA dependency at the supply-chain level. 4. 512K context + hybrid thinking in a 2B model, with native tool calling, deep search, and code generation — a foundation for on-device coding agents. 5. UltraX operationalizes "row-level add/delete/edit" — data governance as programs, not cleaning.

    Risks and Open Questions

  • Open-source date undetermined — the announcement only says "soon", which matters for community ecosystem growth.
  • AA-Index is a composite — the 54.26 average spans 7–8 benchmarks; single-domain extremes (math/code) may not lead; domain-specific boards need checking.
  • τ²-Bench is a single scenario — the 90.35 in telecom customer service cannot be extrapolated to other agent scenarios (booking, office workflows, etc.) without broader agent benchmarks.
  • True fidelity of the 9-chip ports — "aligned with the NVIDIA platform overall" is claimed, but boundary conditions of "some scenarios match or exceed" were not disclosed.
  • One-Line Takeaway

    A 2B model beating an 8B sparse model, τ²-Bench Telecom 90.35, and Day-0 adaptation on 9 domestic chips: ModelBest has pushed the on-device agent base model to the lowest price point while reducing NVIDIA dependency — formally leaving Qwen3.5 and Gemma 4 behind in the sub-4B class.

    Source links:

  • Original WeChat article: https://mp.weixin.qq.com/s/rjFxrUylyGMqa5QtgypCdw
  • Jiazi Guangnian: https://so.html5.qq.com/page/real/search_news?docid=70000021_4606a5cc9aa37852

Tags

#minicpm5-2b#modelbest#on-device-ai#agent-models#edge-computing#benchmarks#domestic-chips#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178446942