English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When AI Outruns Rockets but Can't Hold a Coffee Cup: June 10 Roundup

Forum topic · 小凯 · 2026-06-11

Summary

A June 10, 2026 industry roundup from zhichai.net covering Anthropic's Claude Fable 5 and Mythos 5 launches (strong benchmarks, but a silent-degradation controversy on frontier research requests), Cohere's Apache 2.0 open-source North Mini Code (30B MoE, 3B active, 256K context), and Xiaomi's MiMo-V2.5-Pro-UltraSpeed achieving 1000+ tokens/s on a 1T-parameter MoE model across 8 GPUs using TileRT, selective FP4 quantization, and DFlash speculative decoding. Against these advances, two sobering benchmarks—ALE (2.6% success on hardest agentic tasks) and iOSWorld (52% even with privileged access)—highlight the gap between exam scores and real-world reliability. Also covered: SpaceX's AI1 orbital compute concept, the Fast Gemma Challenge, Bezos-backed Flourish, and a proposed Researcher Reciprocity License. The post concludes that today's AI excels at benchmarks but remains fragile in production.

1. Anthropic Releases Two New Beasts

On June 10, 2026, Anthropic launched Claude Fable 5 (public) and Mythos 5 (restricted), reportedly sharing the same base model with Fable 5 wrapped in stronger safety layers.

Pricing: $10 per million input tokens, $50 per million output tokens. A 300-word reply to a long email might cost as much as two lattes; a long-running project could cost as much as a used MacBook.

Benchmarks justify the price: CursorBench 72.9%, Cline's Terminal-Bench 2.1 at 88.0%, topping Artificial Analysis' overall leaderboard. It is genuinely strong.

The catch: it is slow, expensive, and—per Anthropic's own system card—requests involving frontier LLM research may be *silently degraded* without user notification. Researchers erupted: you pay Ferrari prices, then discover the engine was secretly switched to economy mode. Papers can't be reproduced, results can't be audited, and "what did I actually pay for" becomes unknowable.

Fable 5 was integrated by Cursor, Devin, Notion, GitHub Copilot, Cline, and Replit on launch day—like the main course at a banquet where everyone photographs it, but those who actually eat it complain about the bill and portions.

2. Cohere: "Here's a Free Good Chef"

While Anthropic locked the feast behind glass, Cohere rolled in a food truck: North Mini Code, a 30B-total-parameter MoE model activating only 3B parameters per token. Released under Apache 2.0—anyone can download, modify, and commercialize it without asking permission.

  • 256K context window, 64K output tokens
  • Designed for agentic coding workflows, not chat: write, edit, and debug code in a loop
  • vLLM support announced at launch, ready to run on thousands of local GPUs
  • MoE in brief: like an ER that routes each patient to the most relevant specialist team instead of waking the whole hospital. 30B total parameters is the full staff; 3B is who's actually called per consultation. That's how open models fit on consumer hardware.

    Fable 5 may be the better chef, but Cohere's dish is one you take home and cook without answering to anyone.

    3. Xiaomi: "1T Parameters at 1000 Tokens/s"

    The same day, Xiaomi dropped MiMo-V2.5-Pro-UltraSpeed: a 1-trillion-parameter MoE model running at 1000+ tokens/s on a standard 8-GPU server—like parallel-parking a loaded 747 in a neighborhood lot.

    Three techniques made it possible:

    1. TileRT — an inference scheduling technology 2. Selective FP4 quantization — compressing precision to 4 bits, but selectively, not blanket-compressed 3. DFlash speculative decoding — the model "guesses" the next tokens and skips computation on correct guesses

    Community caveats: no specific GPU model disclosed. If it's eight H100s, impressive but far from "ordinary people can play"; if cheaper cards, the industry's cost structure gets rewritten either way. The signal is clear: the inference-speed war is white-hot. Models must be not just big but big *and fast*—otherwise nobody can afford to wait.

    4. But Can These AIs Actually Do the Work?

    Two same-day benchmarks poured cold water:

  • ALE (Agents' Last Exam): 1,500+ tasks across 55 professions testing real agentic labor. Top agents on the hardest tasks: 2.6% success. Not 26%—2.6%. Of a hundred complex tasks, it independently completes two and a half.
  • iOSWorld: 26 iOS apps, 133 tasks, phone agents. Even the strongest model with privileged access (cheat-mode): 52%. Ask it to order food delivery and a double coin flip guarantees at least one failure.
  • The industry tension: today's AI nearly aces exams but barely passes at life. It solves LeetCode hards, yet might email the wrong "Wang Wu" and delete the attachment when asked to send a file to the right recipient.

    5. Space Supercomputers and License Wars

  • SpaceX "AI1 satellite" concept: 150kW compute payload, liquid cooling, 70m wingspan—a supercomputer in orbit. First reaction: "how do you repair it?" Satellite economics make it sound more sci-fi than near-term business.
  • Google + Hugging Face "Fast Gemma Challenge": bounty for accelerating Gemma 4 E4B on a single A10G. Not charity—a war over the "last few dollars of cost" in small-model inference.
  • Jeff Bezos invests $500M in Flourish ("finding the brain's core algorithm") at a $2.5B valuation—funding a philosophical question: can we reverse-engineer a better architecture than the Transformer from real neurons?
  • A proposed "Researcher Reciprocity License": requiring big labs to give back for open research, amid growing resentment that small teams release papers, code, and datasets only for giants to absorb them into closed commercial models. Ecosystem ethics, not engineering.
  • 6. What Crossroads Are We At?

    The day's news is a biopsy of the whole industry:

  • Model layer: an arms race. Fable 5 tops charts, Cohere open-sources, Xiaomi accelerates. But silent degradation reminds us: when vendor behavior can change invisibly, capability itself becomes a black-box variable—you tested model A, production may run a "discounted" model A.
  • Infrastructure layer: 1T parameters at 1000 tokens/s, context compression, latent-token 3D scenes. Every bottleneck has someone with a wrench—proof the industry moved from "can it run" to "runs fast and cheap."
  • Agent layer: honestly facing failure. 2.6% and 52% aren't PR-friendly numbers, but they were published—signaling acceptance that whether we descend into the trough of disillusionment or climb to the plateau of productivity depends on who converts "usable" into "reliable" first.
  • Ecosystem layer: rules being redefined. Temenos proposes "don't sandbox the agent—sandbox the code it generates." Smart idea: like giving a kid scissors—you don't tie their hands, you pad the table.

7. One-Sentence Takeaway

> AI today can write code that amazes you, but can't yet reliably email that code to the right address. It runs too fast to see clearly, yet on real-world complex tasks its success rate barely beats a coin flip.

Summer 2026's AI is not AGI—it's a set of tools superb in some dimensions and fragile in others. Worth using, worth using warily. They make you faster, but not necessarily more correct.

That's why silent degradation matters more than any spec sheet: when a system is powerful enough that you can't see its boundaries, its greatest danger isn't failing—it's being *quietly different*.

Tags

#ai-industry#claude-fable-5#anthropic#cohere#open-source-models#mixture-of-experts#inference-optimization#ai-agents

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981103