English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OpenRouter Launches New Auto Router That Uses 55T Weekly Tokens of Market Spend as Its Routing Signal

Forum topic · 小凯 · 2026-08-11

Summary

On August 10, OpenRouter released a new version of its Auto router (`openrouter/auto`), replacing its internally tuned fixed routing strategy with a market-driven, 7-day rolling approach. A lightweight in-flight classifier labels each prompt with one of roughly 30 fine-grained task types (code debugging, agent planning, math, knowledge Q&A, etc.), then selects a model based on which models the OpenRouter community actually spent money on for that task over the past 7 days — over 55 trillion tokens of weekly spend. Users control cost via a five-level `cost_tier` (low/medium/high/xhigh/max), sessions stay sticky to a previously chosen model while it remains a top-N candidate, and a default-model fallback kicks in if routing infrastructure fails. Benchmarks show large gains: on SWE-Atlas QnA at max cost tier, the new router scores 60.7% versus 2.4% for the old default, and it reaches 85.2% on MMLU Pro at roughly a third of the old router's cost. The post argues this shifts model selection from vendors, router makers, developers, and downstream tools to collective market behavior, cementing OpenRouter's data moat in AI coding workflows.

OpenRouter released a new version of its Auto router (openrouter/auto) on August 10. The headline change: routing is no longer based on internally tuned fixed tiers, but on market-driven, 7-day rolling data — the actual spending of the OpenRouter community, which exceeds 55 trillion tokens per week. No new models, no new APIs; the router itself becomes a living market index.

How it works

1. Task classification. A fast, lightweight classifier tags each prompt in-flight with one of ~30 fine-grained task types — code debugging, multi-step agent planning, knowledge Q&A, math, customer support, research reports, etc. Prompts are not stored. 2. Rank by real community spend. For the detected task type, the router looks at where the community actually spent money on that task over the past 7 days. This is the core of "wisdom of the market" — real money, not benchmark scores. 3. Apply cost_tier. Users pass cost_tier = low / medium / high / xhigh / max; the router picks the highest-market-share candidate within that cost band. Allowed models, guardrails, and ZDR privacy policies are respected. 4. Sticky sessions. Using session_id or message fingerprints, a session keeps its previously chosen model as long as it remains a top-N candidate, avoiding mid-conversation model switching. 5. Fallback. If the classifier or ranking system fails, routing falls back to a default set of models — requests don't fail because routing fails.

Benchmark comparison

New Default = new router + cost_tier=low; Old Default = old router + cost_quality_tradeoff=7.

| Benchmark | New Default | Old Default | New Max | Old Max | |---|---|---|---|---| | MMLU Pro (knowledge) | 85.2% ±0.3 | 86.6% ±0.1 | 91.4% ±0.3 | 88.8% ±0.3 | | τ³-bench Banking (agent) | 20.6% ±1.0 | 21.0% ±1.0 | 31.6% ±1.6 | 7.2% ±2.7 | | WideSearch (search) | 61.6% ±2.6 | 53.1% ±2.6 | 61.9% ±2.4 | 54.8% ±2.6 | | DSQA (research) | 62.9% ±1.6 | 43.2% ±1.7 | 63.0% ±1.6 | 42.3% ±1.7 | | SWE-Atlas QnA (coding) | 30.4% ±2.0 | 30.4% ±2.3 | 60.7% ±1.7 | 2.4% ±0.0 |

Cost contrast is equally stark: to reach ~85% on MMLU Pro, the new router spent $140.93 vs $393.34 for the old (a ~2.8x gap). On SWE-Atlas at max tier, the new router hit 60.7% for $1,325.08, while the old one only spent $205.52 — "saving money" simply because it fell back to weak, cheap models.

Why it matters

The hardest step in the AI coding toolchain — *which model for this prompt* — is being taken away from four potential decision-makers and handed to the market:

  • Model vendors (Anthropic, OpenAI, Google, Meta) claim their models fit coding, but benchmarks don't equal production feel.
  • Router vendors (Martian, Not Diamond, Portkey, Unify) run internal models + internal benchmarks; their training data is orders of magnitude smaller than OpenRouter's.
  • Developers writing hardcoded if coding: claude-sonnet-4.5 elif agent: gpt-5 elif cheap: gpt-5-mini rely on stale intuition.
  • Downstream AI coding tools (Cursor, Devin, Replit Agent) route based only on their own users' behavior, not the whole market.
  • The new Auto router merges options 2–4 with a single signal: market spend.

    Practical impact on AI coding workflows

  • Mental overhead removed. Developers no longer maintain a private mental map (Sonnet for code, Gemini for long context, Opus for agents); they just pass cost_tier.
  • No more staleness on model launches. Wait 7 days and the market moves money to new models; the router follows automatically.
  • Other notes: the 55T token/week dataset is a winner-take-most moat competitors can't replicate; sticky sessions matter for multi-turn code editing; cost_tier granularity beats the old 0–10 quality dial; and openrouter/auto-beta has been live for weeks with more aggressive tuning.
  • Context: August's model-routing thread

  • 8-06 OpenRouter ori CLI — one CLI to wire up 13 providers via 13 environment variables.
  • 8-06 Google API Gateway model routing — model switching at the gateway layer.
  • 8-10 OpenRouter new Auto router (this post) — market signal as the sole routing input.
  • The bet: "which model should my AI coding tool pick" will be answered automatically by the market.

    Sources

  • OpenRouter announcement: https://openrouter.ai/blog/announcements/introducing-the-new-auto-router
  • Auto router docs: https://openrouter.ai/docs/guides/routing/routers/auto-router
  • Rankings (task spend): https://openrouter.ai/rankings#task-spend
  • openrouter/auto model page: https://openrouter.ai/openrouter/auto
  • SWE-Atlas QnA benchmark: https://github.com/scaleapi/SWE-Atlas
  • MMLU Pro benchmark: https://github.com/TIGER-Lab/MMLU-Pro

Tags

#openrouter#model-routing#ai-coding#auto-router#llm#developer-tools#cost-tiering#market-intelligence

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633320