English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Microsoft Switches GitHub Copilot's Default Engine to Its Own MAI Models in August, Trained From Scratch on Maia 200 Chips

Forum topic · 小凯 · 2026-08-13

Summary

According to an August 2026 report circulating on zhichai.net, Microsoft has begun routing production traffic for Excel and Outlook to its in-house MAI model family and switched GitHub Copilot's default backend from GPT-4 Turbo to MAI-Code-1-Flash. The flagship MAI-Thinking-1 reasoning model — a roughly 1-trillion-parameter MoE with 35B active parameters — reportedly matches Claude Opus 4.6 on SWE-Bench Pro (~52.8%), scores 97.0% on AIME 2025, and beats Claude Sonnet 4.6 in a 1,276-task Surge human blind test. Mustafa Suleyman's team emphasizes the models were trained entirely from scratch without distilling outputs from OpenAI, Anthropic, or Google, giving Microsoft full ownership of its data and RLHF pipelines. MAI-Thinking-1 runs on Microsoft's in-house Maia 200 chip, claimed to deliver 1.4x performance-per-watt and 30% better price-performance versus NVIDIA GB200. MAI-Code-1-Flash is priced at $0.75/$4.50 per 1M tokens, over 50% cheaper than GPT-4 Turbo, and leads SWE-Bench Pro at 51.2%. All benchmarks are Microsoft self-reported, and pricing for MAI-Thinking-1 remains undisclosed.

Microsoft's AI division under Mustafa Suleyman has moved from announcements to production: as of August 2026, Bloomberg reports that thousands of prompts in Excel and Outlook are being routed to Microsoft's in-house MAI models, and GitHub Copilot's default backend has switched from GPT-4 Turbo to MAI-Code-1-Flash. MAI-Thinking-1, the family's flagship reasoning model, is in private preview on Azure Foundry with public preview expected within weeks.

MAI-Thinking-1 Benchmark Highlights

| Dimension | MAI-Thinking-1 | Notes | |---|---|---| | Total / active params | ~1T / 35B MoE | — | | Context window | 256K | — | | AIME 2025 | 97.0% | — | | AIME 2026 | 94.5% | — | | SWE-Bench Pro | ~52.8% | Roughly on par with Claude Opus 4.6 | | Terminal Bench 2 | 54.8 | — | | Surge blind test | Beats Claude Sonnet 4.6 overall quality | 1,276 tasks, human raters blind to model identity | | Pricing | Undisclosed | — |

Parity with Claude Opus 4.6 on SWE-Bench Pro — a full-repository, multi-file-edit benchmark — is the key result, suggesting Microsoft now has a credible enterprise coding backend. The Surge blind test is notable because independent human raters compared outputs without knowing which model produced them, arguably closer to real-world perceived quality than static benchmarks like GSM8K or MMLU.

Why "No Distillation" Matters

Mustafa Suleyman stated at Build 2026 that the MAI family was "trained from the ground up, without knowledge distillation from any third-party models." Most reasoning models today are distilled students of closed-source flagships (OpenAI o1/GPT-4o, Anthropic Sonnet/Opus, Gemini), which caps their ceiling at the teacher's ability.

Two engineering implications:

1. Full pipeline ownership — Microsoft owns its entire SFT/RLHF data pipeline, its preference labeling, and its reward models, so the resulting "model taste" is its own. 2. Independent ceiling — Microsoft can push capability forward on its own schedule rather than waiting for the next OpenAI release to distill.

Maia 200 Chip: 1.4x Perf/Watt, 30% Price-Performance Edge

MAI-Thinking-1 runs on Microsoft's own Maia 200 chip, not NVIDIA's GB200. Reported figures:

  • MAI-Thinking-1 on Maia 200 vs. the same model on GB200: 30% advantage in performance-per-dollar
  • MAI models end-to-end on Maia 200: 1.4x performance-per-watt
  • Combined with non-distilled training and commercially licensed, traceable data, this forms a full-stack closed loop: Microsoft trains its own models, runs them on its own silicon, and sets its own pricing — a supply-chain answer to OpenAI API dependency.

    Copilot's August Default Switch

    Two Bloomberg-confirmed developments:

    1. Excel and Outlook already route thousands of production prompts to MAI. 2. GitHub Copilot's default engine switched to MAI-Code-1-Flash in August 2026; GPT-4 Turbo remains as fallback until November 2026.

    MAI-Code-1-Flash went GA on June 26, priced at $0.75/1M input and $4.50/1M output tokens — over 50% cheaper than GPT-4 Turbo. It scores 51.2% on SWE-Bench Pro (16 points ahead of Claude Haiku 4.5 at 35.2) and can use up to 60% fewer tokens. VS Code has switched in parallel.

    Enterprise-Grade Training Data as a Differentiator

    Suleyman describes MAI's training data as:

    > "Training data is enterprise-grade, commercially licensed, and traceable — a critical distinction for organizations with compliance requirements."

    This matters for GDPR, HIPAA, SEC, and PHI compliance scenarios where OpenAI and Anthropic's data provenance is not fully disclosed.

    Timeline

  • June 2026 (Build): MAI-Thinking-1 plus six-model MAI family (MAI-Image-2.5, MAI-Voice-2, MAI-Transcribe-1.5, MAI-Code-1-Flash, etc.) announced
  • Aug 9: OpenAI pauses part of Astra development (math + cybersecurity)
  • Aug 10: OpenAI releases GPT-5.6-Cyber
  • Aug 12: MAI-Thinking-1 launch publicized via Suleyman's X account
  • August: Bloomberg confirms Excel/Outlook traffic routing and Copilot default switch
  • Caveats

  • MAI-Thinking-1 is still in private preview; timing of public availability is uncertain
  • Pricing is undisclosed
  • All benchmarks are Microsoft self-reported; independent leaderboards have not yet updated
  • The "no distillation" claim lacks a published list of training data sources
  • Maia 200 comparisons don't specify hardware configs or inference frameworks vs. GB200
  • OpenAI still runs on Azure under existing commercial agreements; MAI is a hedge, not a rip-and-replace
  • Strategic Read

    Microsoft has invested over $13B in OpenAI, yet its real move here is cultivating a second hand rather than replacing its partner. Keeping OpenAI models as Copilot fallbacks while defaulting to MAI signals negotiating leverage: when OpenAI's prices rise or access tightens, Microsoft can switch seamlessly. This is vendor-level "sovereign AI" — echoed by the Mayo Clinic partnership for co-created, customer-owned vertical models — and it positions MAI not as an OpenAI API replacement but as a platform for enterprise-owned AI deployments via Microsoft Foundry.

    Sources

  • https://x.com/mustafasuleyman
  • https://www.bighatgroup.com/blog/microsoft-ai-weekly-2026-08-02
  • https://byteiota.com/microsoft-mai-model-family-four-new-models-at-build-2026/
  • https://www.ndtvprofit.com/markets/bye-bye-copilot-microsoft-unveils-its-own-proprietary-new-ai-models-11583588
  • https://www.techmeme.com/260602/p47
  • https://m.marsbit.co/flashshare/20260603084506413513.html
  • https://www.theverge.com

Tags

#microsoft#github-copilot#mai-thinking-1#maia-200#maicode-1-flash#mustafa-suleyman#openai#inference-chips

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633412