Overview
MiniMax M3 is presented as the first Chinese flagship model to combine three capabilities simultaneously: a 1M-token context window, native multimodality, and frontier coding & agentic performance — with API prices significantly below overseas competitors.
Technical Foundation
MSA Sparse Attention
| Metric | Data | |--------|------| | Context window | Up to 1M tokens, with at least 512K guaranteed | | Efficiency gain | Per-token computation at 1M scale reduced to ~1/20 of the previous generation | | Inference optimization | Redesigned data reading and compute paths, 4x+ performance |
Traditional full attention has exploding compute costs at long context; MSA's sparsification brings it to a practical level.
Native Multimodality
- Trained from step zero on mixed text, image, and video data
- Pretraining data expanded to the hundreds-of-TBs scale
- Highly aligned text and vision semantic spaces
- Supports image understanding, video understanding, and desktop operation (Computer Use)
- Runtime: nearly 12 hours
- Autonomous output: 18 commits + 23 experiment charts
- Result: core experiments ran successfully
- Score: 37.1 (third place)
- Ranking: Opus 4.7 (42.4) > GPT-5.5 (39.3) > M3 (37.1)
- 2026-05-29: IPO coaching agreement with CITIC Securities for an A-share listing
- Hong Kong market cap has surpassed HK$236 billion
- M3's release is a pre-listing showcase of technical strength
- Past four years: leading models built separately for text, speech, video, and music
- From 2026: pushing deep cross-modal fusion
- M3 is the text flagship; the Hailuo 3.0 video model is coming soon
Performance
Benchmarks
| Test | MiniMax M3 | Comparison | |------|-----------|------------| | BrowseComp | 83.5 | Opus 4.7: 79.3 | | SWE-Bench Pro | Surpasses GPT-5.5 & Gemini 3.1 Pro | Approaches Opus 4.7 | | SVG-Bench | Surpasses Opus 4.7 | — |
Case Study: Reproducing an ICLR 2025 Paper
MiniMax gave M3 the ICLR outstanding paper "Learning Dynamics of LLM Finetuning" and had it independently reproduce the work:
Capability combination: multimodal understanding of charts and formulas + long context fitting papers/code/logs + coding/agent skills driving long-horizon execution.
Autonomous Model Training
Given 4 base models with only pretraining completed, M3 autonomously completed data synthesis, training, evaluation, and iteration within 12 hours:
No human intervention throughout.
Pricing
API (context ≤512K, limited-time 50% off)
| Item | Standard | Priority | |------|----------|----------| | Input | 2.1 CNY / M tokens | 3.15 CNY | | Output | 8.4 CNY / M tokens | 12.6 CNY | | Cache read | 0.42 CNY / M tokens | 0.63 CNY |
Token Plan Subscriptions
| Tier | Monthly fee | Credits | |------|-------------|---------| | Plus | 49 CNY | 600M tokens | | Max | 119 CNY | 1.8B tokens | | Ultra | 469 CNY | 5.5B tokens |
Features: reduced deduction for simple tasks, a unified credit pool, and notably more tokens per price point than overseas equivalents.
Strategic Context
Ahead of IPO
Omni-Modal Strategy
Token Economics Bet
CEO Yan Junjie predicts that AI coding, office work, and multimodal use will drive 10-100x growth in token consumption in 2026. M3's MSA architecture is designed for that growth.
Limitations
1. Exact benchmark scores not fully disclosed (SWE-Bench Pro "approaches Opus 4.7" but no exact value given) 2. "At least 512K guaranteed" suggests 1M may be a ceiling rather than a stable value; long-context attention quality degradation not disclosed 3. "Open-sourcing soon" timing uncertain; community maintenance to be observed 4. Fewer ecosystem tools and third-party integrations; enterprise-grade feature maturity unknown 5. IPO pressure: pre-listing launches may face pressure to "hit deadlines"
Competitive Landscape
| Model | Context | Multimodal | Coding | Cost | |-------|---------|-----------|--------|------| | MiniMax M3 | 1M | Native | Near Opus 4.7 | Low | | Claude Opus 4.7 | 200K | Limited | SOTA | High | | GPT-5.5 | 128K | Supported | Strong | High | | Gemini 3.1 Pro | 1M | Native | Strong | Mid-high |
MiniMax's differentiation: at the intersection of 1M context + native multimodality + strong coding, it is significantly cheaper.
Key Takeaways
1. Technical breakthrough: MSA sparse attention makes 1M context practical, cutting computation to 1/20 2. Capability combination: first Chinese model uniting coding + 1M context + native multimodality 3. Cost advantage: API pricing and Token Plans well below overseas equivalents 4. Validation: 12-hour autonomous paper reproduction with 18 commits and 23 charts demonstrates long-horizon execution 5. Strategic moment: a key pre-IPO release and a milestone in the omni-modal strategy 6. Risks: undisclosed exact scores; uncertain open-source timing; immature ecosystem
---
Sources: MiniMax official, IT Home, DoNews, Sina Finance, Phoenix Finance. Research date: 2026-06-05.