Meta Muse Glimmer 30B: The First Open-Weight Model Truly Designed for Local Agentic Workflows, Bringing the RTX 5090 Back into the AI Coding Toolchain
On August 10, Meta Superintelligence Labs, together with Scale AI, released Muse Glimmer — a 30B-parameter multimodal dense model purpose-built for "local, always-running agentic workflows," under Apache 2.0 open weights. On the same day, LMSYS's SGLang team added it as a day-0 supported model, lowering the deployment floor to "a single RTX 5090." Its biggest difference from previous open-weight large models: it was not designed to top benchmarks, but to run agent loops on your desktop 24/7.
Architecture: Unusual Decisions on the Table
- 30B total = 27.9B text decoder + 1.9B ViT + a GELU multimodal projector. The text decoder is a 52-layer transformer, each layer using grouped-query attention (32 query heads / 2 KV heads) + SwiGLU FFN. The vision side is an independent 1.9B-parameter ViT — it does not eat into the text side's capacity.
- Hybrid attention pattern — three 2048-token sliding-window layers alternating with one full-sequence attention layer. Local windows use RoPE; full-attention layers use NoPE. The combination lets the model push context length beyond the training limit.
- 128k+ token context window — for agent workflows, that means fitting 4–5 full repository codebases plus entire MCP tool protocols plus multi-turn conversation history, with no summarization or compression needed.
- Trained for agents without benchmark regression — per LMSYS figures, compared to leading same-size models, Muse Glimmer remains "excellent" on key agentic use cases and benchmarks (detailed numbers await Meta's paper).
- Performance ceiling — RTX 5090 + NVFP4 + DFlash speculative decoding delivers 236 tok/s/user (batch 1) and 1452 tok/s total throughput (batch 8) on a single card. DFlash accelerates 1.9x–4.3x (236 / 63.9 ≈ 3.7x on the RTX 5090). These numbers flatten the entry barrier for "lightweight agent backends" running 7B–30B models on a single consumer GPU.
- Multiple hardware paths — NVFP4 via the NVIDIA SM120 backend (RTX 5090 / RTX PRO 6000 / DGX Spark); BF16 on a single H100; GGUF Q4KM (18 GB) and Q4K-Dynamic within a 24GB VRAM ceiling; MLX on Apple Silicon Mac mini / M-series MacBook Pro. Developers can migrate seamlessly across MacBook / RTX 5090 / H100 tiers with the same weights.
- Three native optimizations — a low-overhead scheduler, RadixAttention prefix cache, and interruptible CUDA graphs. The prefix cache is critical for agent scenarios: multi-turn conversations, repeated tool schema calls, repeated code reviews — the prefix cache hits these "repeated prefixes" at zero cost.
- DFlash speculative decoding — separates the draft model from the target model; the target only validates multi-token candidates proposed by the draft, roughly doubling throughput without quality loss. It marks a new inflection point in the "latency vs. cost" curve for agent workloads.
- The past 12 months — Anthropic Claude Sonnet / OpenAI GPT-5 / Google Gemini 2.5 Pro were the "cloud default," and locally you only had "barely usable" 7B–13B options (Qwen2.5-Coder-7B, Code Llama 7B, DeepSeek Coder 6.7B). Running agents locally was nearly fantasy.
- The past 6 months — the open-source camp gained 32B-class agent models (Qwen3.5-32B, DeepSeek-V3-Flash, GLM-4.5-Air), but none were designed for "24/7 always-on" operation. They are "benchmark models" — inference memory peaks, long-context KV caches, frequent tool schema calls, and prefix cache hit rates were never optimized for real agent usage.
- Muse Glimmer's answer — 30B total parameters + 128k+ context + runs in 24GB VRAM + Apache 2.0 + SGLang day-0 optimizations + DFlash + RadixAttention prefix cache + a cross-hardware MLX path — the first open-weight model "born for local agents."
- "Agent backend on your own machine" becomes feasible for the first time — previously, 70B models needed 2–4 H100s or 8–16 RTX 4090s, impossible locally on 24GB; 7B models weren't smart enough. Muse Glimmer hits 236 tok/s on an RTX 5090 with 30B + NVFP4 + DFlash, meaning harnesses like Claude Code / Codex / OpenCode finally have a backend option that actually runs locally and is smart enough.
- Privacy/compliance scenarios open up — legal, medical, financial, government, and consulting firms with "code cannot leave company machines" requirements previously had to settle for 7B models or build private data centers. Muse Glimmer extends this path to "a single RTX 5090 workstation."
- The cost structure of always-on agents changes — Claude Sonnet API pricing at $3/$15 per 1M tokens means an always-on agent burning 200k tokens overnight costs $3; running Muse Glimmer locally overnight costs only electricity (RTX 5090 ~400W). Developers can run "re-run the entire test suite every night" style long jobs without watching API bills.
- Scale AI co-authorship — Alexandr Wang confirmed on August 10 that Muse Glimmer 30B + Muse Spark 1.2 were open-sourced simultaneously. Scale provided the data / RLHF pipeline; Meta Superintelligence Labs provided model architecture / training. This marks Scale AI's evolution from "data labeling company" to "co-signing model builder" — its moat in RLHF training pipelines is now front and center.
- Apache 2.0 is the key — Meta's earlier Llama family used the "Llama Community License," with commercial restrictions for products serving 700M+ monthly active users. Apache 2.0 removes all such limits. An underrated detail: for AI coding tool vendors (Cursor / Devin / Replit Agent / OpenChamber), the "license risk" of local backends is eliminated for the first time.
- DFlash + RadixAttention is SGLang's moat — vLLM can also serve 30B models, but SGLang's integrated combination of prefix caching + speculative decoding + multi-hardware paths is what lets it capture the dividends of this 30B era.
- Apple Silicon MLX path — Muse Glimmer achieves 56.9 tok/s on an M5 Pro (Q4K-Dynamic, batch 1), meaning MacBook Pro users can run a 30B agent locally — previously this lane only had make-do 7B models.
- Aug 4 — Ant Ling-3.0-flash open-sourced (topicId 178603081) — end-to-end domestic open stack (KDA+MLA + 1/64 sparse MoE + SGLang HiCache + Mooncake).
- Aug 5 — NVIDIA Cosmos 3 (topicId 178603066) — embodied world models in 64B / 16B / 4B tiers.
- Aug 9 — Microsoft SkillOpt (topicId 178603080) — portable experience layer.
- Aug 10 — Meta Muse Glimmer 30B (this article) — the first open-weight model designed for always-on local agents.
- LMSYS SGLang blog: https://www.lmsys.org/blog/2026-08-10-meta-muse-glimmer
- Meta @AIatMeta announcement: https://x.com/AIatMeta/status/2086757844544811485
- Scale AI @alexandr_wang announcement: https://x.com/alexandr_wang/status/2086756152034066792
- Meta Research details: https://research.meta.ai/blog/introducing-muse-glimmer-open-agentic-model
- SGLang installation: https://github.com/sgl-project/sglang
- DFlash speculative decoding: https://github.com/sgl-project/sglang/tree/main/python/sglang/speculative
SGLang Day-0 Support: From "Released" to "Ready Out of the Box"
LMSYS's blog lists very concrete capability details:
Reframing the Local Backend for AI Coding
What This Means for AI Coding Workflows
Notable Details
Context: The August AI Coding Toolchain Timeline
August 10 is the moment the "localization era of AI coding" formally begins — developers finally have a genuinely comparable choice between "cloud Sonnet/Opus" and "local 30B," rather than "cloud vs. a make-do 7B."