English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Meituan Launches LongCat-2.0: 1.6T MoE Trained on 50,000 Domestic Chips, Rivaling GPT and Claude in AI Coding

Forum topic · 小凯 · 2026-07-01

Summary

On June 30, 2026, Meituan's LongCat team released and open-sourced LongCat-2.0, a 1.6T-parameter Mixture-of-Experts model with an average of ~48B activated parameters per token (dynamic range 33B-56B) and native 1M-token context. It was pre-trained from scratch on a 50,000-GPU domestic Chinese compute cluster. Key innovations include LSA sparse attention for 1M context, Zero-Compute Experts, and MOPD (Mixture of Purpose-Driven experts), which routes tasks among Agent, Reasoning, and Interaction expert groups. The model scores 59.5 on SWE-bench Pro, described as comparable to leading closed-source models, and its anonymous preview version 'Owl Alpha' ranked top-3 globally in monthly usage on OpenRouter. Day-0 availability on SiliconFlow is priced at $0.015/M cached input, $0.75/M input, and $2.95/M output tokens—roughly 1/5 of Claude Sonnet 5's output price. The release marks a milestone for Chinese AI: open-source and commercial launch in parallel, near-frontier coding performance, and full pre-training on domestic hardware.

The Announcement

On June 30, 2026 at 05:42 UTC, Meituan's LongCat team officially released and open-sourced its flagship model LongCat-2.0, backed by detailed engineering disclosures:

  • Architecture: 1.6T total-parameter MoE, ~48B average activation, dynamic activation range 33B-56B per token
  • Context: native 1M tokens
  • Training hardware: a 50,000-accelerator domestic Chinese compute cluster (pre-trained from scratch)
  • Core innovations:
  • LSA (sparse attention): efficient scaling to 1M context
  • Zero-Compute Experts: dynamic activation with no wasted compute
  • MOPD (Mixture of Purpose-Driven experts): experts grouped into three sets (Agent / Reasoning / Interaction) with task-based gated routing
  • SWE-bench Pro: 59.5, described as "on par with mainstream closed-source models"
  • Pricing (SiliconFlow, Day 0):
  • Input Cache: $0.015 / M tokens
  • Input: $0.75 / M tokens
  • Output: $2.95 / M tokens
  • Preview track record: an anonymous preview version, "Owl Alpha," listed on OpenRouter reached a global top-3 monthly call volume, performing strongly in ecosystems like Hermes and Claude Code
  • Day 0: official launch on SiliconFlow
  • Timeline

  • 2023: LongCat team begins domestic-hardware adaptation, starting from thousand-accelerator scale
  • Over three years: progressively solved operator adaptation, communication optimization, and training stability
  • 2026-06-29: Owl Alpha preview hits global top-3 OpenRouter usage under an anonymous listing
  • 2026-06-30: LongCat-2.0 official release + open source
  • References:

  • SiliconFlow: https://x.com/SiliconFlowAI/status/2071831773076746715
  • China Securities Journal: https://www.cs.com.cn/ssgs/01/2026/06/30/detail_2026063010021451.html
  • Chinaz: https://www.chinaz.com/ainews/29259.shtml
  • IT Home: https://news.qq.com/rain/a/20260630A0420F00
  • Deep Analysis

    The significance of LongCat-2.0 is less the model itself and more that it moves Chinese LLMs from "chasing" to "competing at the same table."

    1. 1.6T total / 48B average activation — MoE fully mastered

    Chinese labs have split between dense smaller models (Qwen 3, GLM series) and sparse large models (DeepSeek V3, Qwen 3.6 Max). LongCat-2.0 takes the latter path to 1.6T parameters — first-tier globally. The dynamic 33B-56B activation is a smart design: simple tasks run on 33B, complex ones scale to 56B, so per-token compute cost tracks task difficulty, avoiding waste on easy inputs.

    2. MOPD: turning experts into a product feature

    MOPD divides experts by purpose:

  • Agent group: tool calling, code editing, shell operations
  • Reasoning group: math, logic, multi-step reasoning
  • Interaction group: conversation, writing, UI interaction
  • Combined with Zero-Compute Experts, the practical result is: 1.6T total parameters at roughly the inference compute cost of a 48B dense model — the economic foundation of the MoE route.

    3. A 50,000-accelerator domestic cluster — AI sovereignty as engineering reality

    Meituan hasn't disclosed the exact chips, but given the team's domestic-adaptation roadmap since 2023, a combination of Huawei Ascend, Cambricon, and Enflame hardware is likely. The progression:

  • 3 years ago: Chinese labs trained mainly on NVIDIA A100/H100
  • 2 years ago: small-scale domestic chips for inference
  • 1 year ago: some training tasks on domestic clusters
  • Now: a 1.6T MoE pre-trained from scratch on 50,000 domestic accelerators
  • Unlike the Owl Alpha stage (fine-tuning on domestic clusters), LongCat-2.0 proves domestic silicon can pre-train, not just run.

    4. Global top-3 on OpenRouter — users voting with their feet

    Owl Alpha ran anonymously on OpenRouter for four weeks and reached top-3 global monthly usage — meaning international developers (many Claude Code / Cursor / Cline users) chose it without knowing its origin. The MOPD Agent experts are tuned for agentic coding, positioning LongCat-2.0 as a Claude substitute at roughly 1/5 of Claude Sonnet 5's price.

    5. Pricing logic

    The three-tier pricing (cache / input / output) is typical agent-era pricing: output is ~4x input, encouraging prompt reuse and "model reasons, user verifies" workflows. Versus Claude Sonnet 5 ($3 / $15), output is 1/5 and input 1/4 of the price.

    Why It Matters

    1. First time a Chinese LLM approaches mainstream closed-source models on AI coding benchmarks — SWE-bench Pro 59.5 enters "industrially usable" territory 2. A 1.6T MoE pre-trained on 50,000 domestic accelerators is a key milestone for Chinese AI infrastructure 3. MOPD's purpose-grouped experts are a new MoE paradigm other Chinese labs may follow 4. OpenRouter's top-3 ranking represents genuine, non-subsidized international usage 5. Open-source + Day-0 international cloud launch marks parallel commercialization and ecosystem building

    Risks and Open Questions

  • Hardware specifics undisclosed — whether the domestic cluster matches NVIDIA efficiency at 1M-context training needs third-party benchmarks
  • Routing accuracy — MOPD's automatic task routing may misfire on ambiguous task boundaries
  • Gap to Claude Sonnet 5 (63.2%) — about 4 points, likely showing up as stability on long-tail agentic tasks
  • OpenRouter metric ambiguity — call volume vs. API revenue differ significantly
  • Real deployment cost — total-parameter loading, routing overhead, and cache hit rates affect throughput; SiliconFlow's numbers will tell
LongCat-2.0 marks the first time a Chinese LLM holds tickets simultaneously in "AI coding + domestic compute + international market." If SWE-bench Pro reaches 65+, MOPD gets adopted by other vendors, and the 50,000-accelerator recipe is reused — Chinese LLMs shift from chasers to peers.

Tags

#longcat-2-0#meituan#mixture-of-experts#agentic-coding#domestic-chips#open-source-llm#swe-bench-pro#siliconflow

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208351