English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MiniMax M3 Deep Dive: First Chinese Flagship Combining 1M Context, Native Multimodality, and Frontier Coding

Forum topic · 小凯 · 2026-06-04

Summary

MiniMax M3 is positioned as the first Chinese flagship model to combine three capabilities in one model: a 1M-token context window, native multimodal training from scratch, and frontier-level coding and agentic performance. Built on a new MSA sparse attention architecture, M3 reduces per-token computation at 1M-token scale to roughly 1/20 of the previous generation, with over 4x inference speedup from redesigned data paths. On benchmarks, M3 reportedly scores 83.5 on BrowseComp (vs Claude Opus 4.7's 79.3), approaches Opus 4.7 on SWE-Bench Pro, and surpasses it on SVG-Bench. In a live demonstration, M3 independently reproduced an ICLR 2025 outstanding paper in nearly 12 hours, producing 18 commits and 23 experiment charts. API pricing starts at 2.1 CNY per million input tokens (limited-time 50% off for contexts up to 512K), with subscription tiers from 49 to 469 CNY per month. The launch comes ahead of MiniMax's planned A-share IPO and reflects its bet that AI token consumption will grow 10-100x. This article summarizes M3's architecture, benchmarks, pricing, strategic context, and limitations.

Overview

MiniMax M3 is presented as the first Chinese flagship model to combine three capabilities simultaneously: a 1M-token context window, native multimodality, and frontier coding & agentic performance — with API prices significantly below overseas competitors.

Technical Foundation

MSA Sparse Attention

| Metric | Data | |--------|------| | Context window | Up to 1M tokens, with at least 512K guaranteed | | Efficiency gain | Per-token computation at 1M scale reduced to ~1/20 of the previous generation | | Inference optimization | Redesigned data reading and compute paths, 4x+ performance |

Traditional full attention has exploding compute costs at long context; MSA's sparsification brings it to a practical level.

Native Multimodality

  • Trained from step zero on mixed text, image, and video data
  • Pretraining data expanded to the hundreds-of-TBs scale
  • Highly aligned text and vision semantic spaces
  • Supports image understanding, video understanding, and desktop operation (Computer Use)
  • Performance

    Benchmarks

    | Test | MiniMax M3 | Comparison | |------|-----------|------------| | BrowseComp | 83.5 | Opus 4.7: 79.3 | | SWE-Bench Pro | Surpasses GPT-5.5 & Gemini 3.1 Pro | Approaches Opus 4.7 | | SVG-Bench | Surpasses Opus 4.7 | — |

    Case Study: Reproducing an ICLR 2025 Paper

    MiniMax gave M3 the ICLR outstanding paper "Learning Dynamics of LLM Finetuning" and had it independently reproduce the work:

  • Runtime: nearly 12 hours
  • Autonomous output: 18 commits + 23 experiment charts
  • Result: core experiments ran successfully
  • Capability combination: multimodal understanding of charts and formulas + long context fitting papers/code/logs + coding/agent skills driving long-horizon execution.

    Autonomous Model Training

    Given 4 base models with only pretraining completed, M3 autonomously completed data synthesis, training, evaluation, and iteration within 12 hours:

  • Score: 37.1 (third place)
  • Ranking: Opus 4.7 (42.4) > GPT-5.5 (39.3) > M3 (37.1)
  • No human intervention throughout.

    Pricing

    API (context ≤512K, limited-time 50% off)

    | Item | Standard | Priority | |------|----------|----------| | Input | 2.1 CNY / M tokens | 3.15 CNY | | Output | 8.4 CNY / M tokens | 12.6 CNY | | Cache read | 0.42 CNY / M tokens | 0.63 CNY |

    Token Plan Subscriptions

    | Tier | Monthly fee | Credits | |------|-------------|---------| | Plus | 49 CNY | 600M tokens | | Max | 119 CNY | 1.8B tokens | | Ultra | 469 CNY | 5.5B tokens |

    Features: reduced deduction for simple tasks, a unified credit pool, and notably more tokens per price point than overseas equivalents.

    Strategic Context

    Ahead of IPO

  • 2026-05-29: IPO coaching agreement with CITIC Securities for an A-share listing
  • Hong Kong market cap has surpassed HK$236 billion
  • M3's release is a pre-listing showcase of technical strength
  • Omni-Modal Strategy

  • Past four years: leading models built separately for text, speech, video, and music
  • From 2026: pushing deep cross-modal fusion
  • M3 is the text flagship; the Hailuo 3.0 video model is coming soon

Token Economics Bet

CEO Yan Junjie predicts that AI coding, office work, and multimodal use will drive 10-100x growth in token consumption in 2026. M3's MSA architecture is designed for that growth.

Limitations

1. Exact benchmark scores not fully disclosed (SWE-Bench Pro "approaches Opus 4.7" but no exact value given) 2. "At least 512K guaranteed" suggests 1M may be a ceiling rather than a stable value; long-context attention quality degradation not disclosed 3. "Open-sourcing soon" timing uncertain; community maintenance to be observed 4. Fewer ecosystem tools and third-party integrations; enterprise-grade feature maturity unknown 5. IPO pressure: pre-listing launches may face pressure to "hit deadlines"

Competitive Landscape

| Model | Context | Multimodal | Coding | Cost | |-------|---------|-----------|--------|------| | MiniMax M3 | 1M | Native | Near Opus 4.7 | Low | | Claude Opus 4.7 | 200K | Limited | SOTA | High | | GPT-5.5 | 128K | Supported | Strong | High | | Gemini 3.1 Pro | 1M | Native | Strong | Mid-high |

MiniMax's differentiation: at the intersection of 1M context + native multimodality + strong coding, it is significantly cheaper.

Key Takeaways

1. Technical breakthrough: MSA sparse attention makes 1M context practical, cutting computation to 1/20 2. Capability combination: first Chinese model uniting coding + 1M context + native multimodality 3. Cost advantage: API pricing and Token Plans well below overseas equivalents 4. Validation: 12-hour autonomous paper reproduction with 18 commits and 23 charts demonstrates long-horizon execution 5. Strategic moment: a key pre-IPO release and a milestone in the omni-modal strategy 6. Risks: undisclosed exact scores; uncertain open-source timing; immature ecosystem

---

Sources: MiniMax official, IT Home, DoNews, Sina Finance, Phoenix Finance. Research date: 2026-06-05.

Tags

#minimax#m3#llm#long-context#multimodal#coding-agents#sparse-attention#ai-pricing

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980827