English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

MiniMax M3 Open-Weights Model Released: 428B Total / 23B Active Params with Coding, Agent, and Long-Context Capabilities

Forum topic · QianXun · 2026-06-13

Summary

On June 12, 2026, MiniMax announced the open-weights release of MiniMax M3 on Hugging Face, positioning it as the first open-weights model to combine three frontier capabilities: coding, agentic task execution, and long-context understanding. The MoE model has approximately 428B total parameters with only 23B active (about 5.4% activation ratio), and uses MiniMax Sparse Attention to scale context length to 1M tokens. Official agentic benchmarks include SWE-Bench Pro at 59.0%, Terminal Bench 2.1 at 66.0%, SWE-fficiency at 34.8%, KernelBench Hard at 28.8%, and MCP Atlas at 74.2%, the latter indicating native support for the Model Context Protocol tool-calling ecosystem. A companion technical report titled 'MiniMax Sparse Attention' is available via Hugging Face Papers (2606.13392). The post notes that license type, training data compliance, and inference cost details remain undisclosed. If licensed permissively, M3 could serve as a foundational model for self-hosted AI coding agents, particularly in compliance-sensitive sectors requiring on-premise deployment and data sovereignty.

On June 12, 2026, MiniMax's official account (@MiniMax_AI) announced that the MiniMax M3 open-weights model is now live on Hugging Face. The company describes it as "The First Open-Weights Model to Combine Three Frontier Capabilities," referring to coding, agentic execution, and long-context understanding.

  • Model repository: huggingface.co/MiniMaxAI/MiniMax-M3
  • Technical paper: huggingface.co/papers/2606.13392 ("MiniMax Sparse Attention")
  • Announcement traction: 379K views, 2,200 likes, 252 reposts at the time of reporting
  • Key Specifications

  • Total parameters: ~428B
  • Active parameters: ~23B (≈5.4% activation ratio)
  • Architecture: Sparse MoE, with activation ratio tighter than the typical 8–15% range, making per-inference cost more controllable
  • Context length: up to 1M tokens, enabled by MiniMax's in-house Sparse Attention mechanism (sliding window + global anchors + cross-layer sparsity), reducing attention compute toward O(n log n) versus the O(n²) cost of standard Transformers
  • Official Agentic Benchmarks

    | Benchmark | Score | What It Measures | |---|---|---| | SWE-Bench Pro | 59.0% | Real GitHub issue auto-fixing | | Terminal Bench 2.1 | 66.0% | Terminal/CLI agent tasks | | SWE-fficiency | 34.8% | Software engineering efficiency (success rate × token economy) | | KernelBench Hard | 28.8% | GPU kernel code generation | | MCP Atlas | 74.2% | Model Context Protocol tool calling |

    For context, the reported SWE-Bench Pro score of 59.0% is slightly above publicly available figures for GPT-5 (~56%) and Claude Code (57%). The 74.2% on MCP Atlas—Anthropic's official MCP evaluation—signals strong native support for tool-calling protocols relevant to the agent ecosystem.

    Why It Matters

    1. Three capabilities in one model: The post frames M3 as marking a shift in open models—from base capability plus long context (Llama 3 era), to reasoning plus coding (DeepSeek V3 era), to the current trifecta of coding + agent + long context, which matches the core bottlenecks of agentic AI in 2026: stable tool calling, multi-step planning, and long-document comprehension.

    2. Sparse attention as the next technical frontier: Alongside DeepSeek NSA, BAAI's HiAttention, and Mistral's sliding-window work, M3's sparse attention extends sparsity from FFN layers to the attention layer itself—a systemic approach to lowering long-context inference cost.

    3. Impact on AI coding toolchains: The Terminal Bench + MCP Atlas combination makes M3 a candidate base model for IDE agents (Cursor, Cline, Roo Code), while open weights enable private, on-premise deployment for compliance-sensitive sectors (finance, government, healthcare).

    Unresolved Questions

  • License type (Apache 2.0, custom, etc.) and commercial-use restrictions are not yet disclosed
  • Training data scale, mix, and compliance—especially for a coding model likely trained on large amounts of GitHub data
  • Inference cost: 428B/23B deployment requires H100/H200-class hardware; per-token pricing will determine practical usability
  • Whether MiniMax will ship an M3-powered IDE or agent platform

Conclusion

MiniMax M3's differentiation lies in the combination of MoE + Sparse Attention + tri-capability (coding/agent/long context) rather than raw leaderboard performance. If the license proves permissive, it could become a de facto base for the open-source AI coding toolchain, and it stands out as one of the most notable open-weights releases of the month for anyone tracking AI coding and agents.

Tags

#minimax#open-weights#moe#sparse-attention#long-context#ai-coding#agents#mcp

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981203