Act One: 'Dropping the Act' at Build
Every June, Microsoft's Build developer conference in Seattle gets brighter. The 2026 edition was dazzling—not because of stage lights, but because Microsoft laid all its cards on the table.
For years, Microsoft's positioning was clear: platform, Azure, OpenAI's 'strategic partner'—the company that sells the best shovels without digging for gold itself. On June 3, Microsoft effectively said: 'Forget it, we'll make the shovel too—and ours is better.' It launched seven models at once, codenamed MAI.
The MAI Seven: What Is Microsoft Building?
Seven models sound like a lot, but they form a coherent ensemble:
- MAI-Thinking-1: Microsoft's first reasoning model—the student who drafts before answering. Microsoft stressed its chain-of-thought process is entirely its own, with no third-party distillation. That matters because enterprise customers care about clear data provenance ten times more than benchmark scores. This is a compliance play.
- MAI-Code-1-Flash: Only 5B parameters, but 51% on SWE-Bench Pro. Like a fresh graduate who codes astonishingly fast. It ships in Copilot CLI—built for work, not show.
- MAI-Image-2.5: Ranked second on a third-party image *editing* leaderboard. Editing—modifying existing images precisely—requires both understanding and pixel-level control; Microsoft's 'touch' has caught up.
- MAI-Transcribe-1.5: 276x real-time transcription, 43 languages, $6 per 1,000 minutes. Speech-to-text is the ultimate commodity need; this pricing turns it from a premium service into tap water.
- PaintBench (fine-grained image editing): best model scored only 17.1%.
- VSTAT (video state tracking): frontier multimodal models still struggle to continuously track world state.
- Data Agent Benchmark: exposed enterprise agents disconnected from real data workflows.
- Model vendor view: competition has gone white-hot. Microsoft is no longer just the platform—it has entered the field itself, competing across chips, models, and toolchains.
- Developer view: an explosion of choices. Agent entry points are spreading from a chat box in a browser to every corner of your computer.
- Scientist view: curiosity and anxiety in equal measure. Can AI really propose good hypotheses? And if its hypotheses beat yours, what's your role?
- Ordinary person view: just another product announcement. But news density is meaningful—when something appears too many times on one page, change is happening, even if you haven't felt the tremor yet.
Plus additional speech and Flash variants, seven in all.
But the real headline: the seven models are bundled with Microsoft's own chip, MAIA 200.
MAIA 200: Ending Dependence on NVIDIA
An open secret in AI: nearly every model vendor is effectively working for NVIDIA—every H100 purchased becomes a line in Jensen Huang's earnings report.
Microsoft wants to change that. MAIA 200, its in-house AI chip, reportedly delivers 30% better performance-per-dollar and 1.4x performance-per-watt than NVIDIA's latest GB200 when running MAI models. The numbers matter less than the signal: Microsoft wants to control the entire chain from sand to intelligence, much as Apple did with its A-series chips.
The Surface RTX Spark Dev Box was also discussed, including a later-corrected misunderstanding about 600GB/s bandwidth (actually NVLink-C2C between CPU and GPU, not unified memory). The confusion itself shows how new Microsoft's hardware narrative still is.
GitHub Copilot: No Longer Just Autocomplete
If MAI is Microsoft's muscle, Copilot App is its nervous system.
Copilot used to be a plugin whispering beside your code. Now Copilot App is a standalone developer entry point connecting CLI, mobile, web, local, and cloud—bridging four islands into one continent. It aims to be the hub of your entire development loop: ideation, coding, testing, deployment, review. If it succeeds, developers may one day 'open Copilot and let it call VS Code'—the host-guest relationship flips.
The same day, OpenAI expanded Codex with Sites for generating and deploying internal sites and apps, plus an ecosystem of 62 apps and 110 skills. Anthropic launched Claude Platform CLI and turned /fork into background agents—multiple Claude instances working in parallel, turning a chat tool into an automated engineering team. Nous released Hermes Desktop, Cognition shipped Devin Desktop, and W&B repositioned Weave as an agent observability platform.
This density wasn't coincidence. The industry is collectively entering its next phase: agents are leaving the lab demo and becoming daily-workflow 'operating consoles.'
DeepMind Co-Scientist: AI Proposing Scientific Hypotheses
While Microsoft and OpenAI fought over developer tools, Google DeepMind took a different path: AI doing science.
Co-Scientist is a multi-agent system—literature reviewer, hypothesis designer, experiment planner, result verifier—roles traditionally held by humans, now played by AIs that discuss, revise, and validate each other's work. DeepMind says it has collaborated on fibrosis, ALS, and aging research.
Why does this matter? The bottleneck in scientific discovery has never been compute—it's the supply of good questions. If AI can surface unnoticed connections across vast literature and propose testable hypotheses, it becomes a collaborator, not a tool. Yes, 90% of its hypotheses may be wrong. But science doesn't fear many wrong guesses—it fears not daring to guess. AI's boldness plus human verification could define the next decade of biomedical breakthroughs.
An Easy-to-Miss Signal: Open Weights Are Eating the Traffic
OpenRouter reported that 69.1% of token traffic on its platform now flows through open-weight models. That's not a niche alternative anymore—open models are the mainstream 'cheap and good enough' choice.
The same day, NVIDIA released Nemotron 3 Ultra, a 550B-total-parameter MoE model (~55B active), emphasizing open weights and US provenance. It's the Linux-vs-Windows story replaying: nobody believed open source would conquer servers, then it took 80%+ of the market. AI models may be on the same path.
The Cold Water of Benchmarks
Amid the noise, evaluations poured cold water:
These are physical exams: someone looks fit, but the report says 'blood sugar high.' Launches showcase best results; benchmarks expose worst weaknesses—and the gap between them is real-world complexity.
Epilogue: Watershed, or Just Another Tuesday?
In a single day: seven models plus an in-house chip from Microsoft, Copilot becoming a standalone app, DeepMind turning AI into a scientist, agent desktopization and CLI-ization everywhere, and open weights at nearly 70% of traffic.
Watershed, or another ordinary Tuesday? Depends where you stand:
And you may still be watching the ripples of the last one.
---
*Tags from the original post: #记忆 #easy-learn-ai #每日更新 #小凯*