Google has quietly shipped model routing in API Gateway (preview & v1), per the August 3 release notes. Teams serving models to coding agents previously had to self-host LiteLLM or hand-roll a proxy; Google now packages that layer into a single serverless gateway URL.
How it works
1. The client sends an OpenAI-compatible request (POST /chat/completions) to the gateway.
2. The gateway reads the model field in the JSON body.
3. It matches routing rules declared via two OpenAPI 3.x extension fields.
4. In flight, it transcodes the request to Vertex AI's prediction schema and dispatches to the Model Garden endpoint; unmatched requests fall back to defaultModel.
Google's own positioning: "a managed alternative to client-side proxies such as LiteLLM."
The documentation examples target three models: google/gemini-3.5-flash-lite, anthropic/claude-opus-4-7, and openai/gpt-oss-120b-maas. Note the third is the open-weight gpt-oss model on MaaS—not OpenAI's closed-source GPT.
What it is not
- Not cross-vendor routing. All three models are hosted on Vertex AI Model Garden; the gateway never connects directly to
api.anthropic.comorapi.openai.com. - No intelligent scheduling. During Preview it "routes based exclusively on the model tag or name in the payload"—static dispatch by model name plus a default fallback, with no cost/latency/quality-based optimization.
- No private/self-hosted models, and all models within a single router must share the same hostname.
- Standard API Gateway pricing: first 2 million calls free; 2M–1B calls at $3.00/million; above 1B at $1.50/million. Model tokens are billed separately via Vertex AI.
- Request timeout cap: 3600 seconds. SSE response streaming is supported.
- Not supported: request-side streaming, gRPC, WebSockets, Gemini Live, VPC Service Controls.
- Known Preview defect: a missing
modelfield in the request body is not surfaced as an error but handled incorrectly. - API Gateway Release Notes (2026-08-03): https://cloud.google.com/api-gateway/docs/release-notes
- Model routing overview: https://docs.cloud.google.com/api-gateway/docs/model-routing-overview
- Configure model routing: https://docs.cloud.google.com/api-gateway/docs/model-routing-configure
- Pricing: https://docs.cloud.google.com/api-gateway/pricing
What it means for AI coding harnesses
The impact is operational, not capability. Enterprises previously ran containers, scaled infrastructure, and managed keys to serve models to agents. Now one edge gateway URL replaces that layer: clients stay unchanged (OpenAI SDK connects directly), while auth, quotas, and usage observability are centralized at the edge. For mixes like "Gemini for long context, Claude for code, gpt-oss for cheap batch work," configuration cost drops noticeably.
It is not a replacement for Apigee AI Gateway—true cross-vendor direct connections, semantic caching, and token budgets remain on Apigee's heavier product line. This is the lightweight version, locked to Google-hosted models.
Pricing and limits
There is no standalone official blog post—only release notes and docs, a documentation-level quiet launch. Chinese reports claiming a chained flow with "Gemini enterprise agent platform, egress via Agent Gateway then API Gateway scheduling" appear in neither official document and should be treated as speculation.