The Announcement
On June 30, 2026 at 05:42 UTC, Meituan's LongCat team officially released and open-sourced its flagship model LongCat-2.0, backed by detailed engineering disclosures:
- Architecture: 1.6T total-parameter MoE, ~48B average activation, dynamic activation range 33B-56B per token
- Context: native 1M tokens
- Training hardware: a 50,000-accelerator domestic Chinese compute cluster (pre-trained from scratch)
- Core innovations:
- LSA (sparse attention): efficient scaling to 1M context
- Zero-Compute Experts: dynamic activation with no wasted compute
- MOPD (Mixture of Purpose-Driven experts): experts grouped into three sets (Agent / Reasoning / Interaction) with task-based gated routing
- SWE-bench Pro: 59.5, described as "on par with mainstream closed-source models"
- Pricing (SiliconFlow, Day 0):
- Input Cache: $0.015 / M tokens
- Input: $0.75 / M tokens
- Output: $2.95 / M tokens
- Preview track record: an anonymous preview version, "Owl Alpha," listed on OpenRouter reached a global top-3 monthly call volume, performing strongly in ecosystems like Hermes and Claude Code
- Day 0: official launch on SiliconFlow
- 2023: LongCat team begins domestic-hardware adaptation, starting from thousand-accelerator scale
- Over three years: progressively solved operator adaptation, communication optimization, and training stability
- 2026-06-29: Owl Alpha preview hits global top-3 OpenRouter usage under an anonymous listing
- 2026-06-30: LongCat-2.0 official release + open source
- SiliconFlow: https://x.com/SiliconFlowAI/status/2071831773076746715
- China Securities Journal: https://www.cs.com.cn/ssgs/01/2026/06/30/detail_2026063010021451.html
- Chinaz: https://www.chinaz.com/ainews/29259.shtml
- IT Home: https://news.qq.com/rain/a/20260630A0420F00
- Agent group: tool calling, code editing, shell operations
- Reasoning group: math, logic, multi-step reasoning
- Interaction group: conversation, writing, UI interaction
- 3 years ago: Chinese labs trained mainly on NVIDIA A100/H100
- 2 years ago: small-scale domestic chips for inference
- 1 year ago: some training tasks on domestic clusters
- Now: a 1.6T MoE pre-trained from scratch on 50,000 domestic accelerators
- Hardware specifics undisclosed — whether the domestic cluster matches NVIDIA efficiency at 1M-context training needs third-party benchmarks
- Routing accuracy — MOPD's automatic task routing may misfire on ambiguous task boundaries
- Gap to Claude Sonnet 5 (63.2%) — about 4 points, likely showing up as stability on long-tail agentic tasks
- OpenRouter metric ambiguity — call volume vs. API revenue differ significantly
- Real deployment cost — total-parameter loading, routing overhead, and cache hit rates affect throughput; SiliconFlow's numbers will tell
Timeline
References:
Deep Analysis
The significance of LongCat-2.0 is less the model itself and more that it moves Chinese LLMs from "chasing" to "competing at the same table."
1. 1.6T total / 48B average activation — MoE fully mastered
Chinese labs have split between dense smaller models (Qwen 3, GLM series) and sparse large models (DeepSeek V3, Qwen 3.6 Max). LongCat-2.0 takes the latter path to 1.6T parameters — first-tier globally. The dynamic 33B-56B activation is a smart design: simple tasks run on 33B, complex ones scale to 56B, so per-token compute cost tracks task difficulty, avoiding waste on easy inputs.
2. MOPD: turning experts into a product feature
MOPD divides experts by purpose:
Combined with Zero-Compute Experts, the practical result is: 1.6T total parameters at roughly the inference compute cost of a 48B dense model — the economic foundation of the MoE route.
3. A 50,000-accelerator domestic cluster — AI sovereignty as engineering reality
Meituan hasn't disclosed the exact chips, but given the team's domestic-adaptation roadmap since 2023, a combination of Huawei Ascend, Cambricon, and Enflame hardware is likely. The progression:
Unlike the Owl Alpha stage (fine-tuning on domestic clusters), LongCat-2.0 proves domestic silicon can pre-train, not just run.
4. Global top-3 on OpenRouter — users voting with their feet
Owl Alpha ran anonymously on OpenRouter for four weeks and reached top-3 global monthly usage — meaning international developers (many Claude Code / Cursor / Cline users) chose it without knowing its origin. The MOPD Agent experts are tuned for agentic coding, positioning LongCat-2.0 as a Claude substitute at roughly 1/5 of Claude Sonnet 5's price.
5. Pricing logic
The three-tier pricing (cache / input / output) is typical agent-era pricing: output is ~4x input, encouraging prompt reuse and "model reasons, user verifies" workflows. Versus Claude Sonnet 5 ($3 / $15), output is 1/5 and input 1/4 of the price.
Why It Matters
1. First time a Chinese LLM approaches mainstream closed-source models on AI coding benchmarks — SWE-bench Pro 59.5 enters "industrially usable" territory 2. A 1.6T MoE pre-trained on 50,000 domestic accelerators is a key milestone for Chinese AI infrastructure 3. MOPD's purpose-grouped experts are a new MoE paradigm other Chinese labs may follow 4. OpenRouter's top-3 ranking represents genuine, non-subsidized international usage 5. Open-source + Day-0 international cloud launch marks parallel commercialization and ecosystem building