Why 72GB Blackwell GPUs Cannot Run DeepSeek V4: SM120 vs SM100 Kernel Gap
Running DeepSeek V4 on consumer Blackwell GPUs like the RTX Pro 5000 (SM120, 72GB GDDR7) fails with an Unsupported architecture error, even though the 144GB combined VRAM of two cards should fit the 148GB model at TP=2 (~73GB per card). The problem is not VRAM; it is kernel support.
Key points
- The hardware fits, the software does not. Two RTX Pro 5000 cards give 144GB GDDR7, which should hold DeepSeek-V4-Flash at TP=2. Model loading succeeds, but inference crashes inside DeepGEMM.
- Blackwell is split into two compute-capability tiers: SM100 (10.0) for datacenter parts such as B200/GB200, and SM120 (12.0) for RTX Pro 5000/6000 and RTX 5090. SM100 and SM120 use different Tensor Core instruction sets; SM100 binaries cannot be relabeled and recompiled for SM120.
- Two V4-specific kernels are missing on SM120.
tf32_hc_prenorm_gemm— used by Manifold-Constrained Hyper-Connections (mHC), a V4-specific layer that is neither standard attention nor standard MoE. SM90 and SM100 implementations exist; SM120 does not.paged_mqa_logits— used by the Lightning Indexer for compressed-attention logits over FP8 KV cache. SM90 and SM100 implementations exist; SM120 does not.- There is no generic CUDA fallback, no Triton path, and no upstream SM120 build.
- Instruction-set mismatch blocks copy-paste fixes. SM100 uses
tcgen05.fence, SM90 useswgmma, and neither exists on SM120. The warp-level GEMM core loop must be rewritten from scratch for SM120's Tensor Core opcodes; changing CMake flags is not enough. - Workarounds, ordered by reliability:
- SGLang + TileLang (partial): mHC layer can be JIT-compiled for SM120, but
paged_mqa_logitsstill has no backend and crashes the attention indexer. - ktransformers community hack: With
SGLANG_DISABLE_DEEP_GEMM=1andSGLANG_OPT_USE_TILELANG_INDEXER=1, RTX PRO 4000 can launch V4, but outputs are all-zero tokens or NaN logits. The torch reference and TileLang fallbacks forfp8_paged_mqa_logitsare numerically wrong on SM120, including a misplacedF.reluthat clips negative attention scores. - Marlin W4A16 fallback (runs but very slow): Disabling DeepGEMM and NVFP4 paths drops throughput to roughly 5 tok/s on a 397B-class model, versus 50+ tok/s for a correct FP4 path. Activation distribution mismatch also breaks the MTP draft head, slowing speculative decoding instead of speeding it up.
- Wait for upstream fixes: Open issues include DeepGEMM #236 (SM120 feature request, Feb 2026), DeepGEMM #317 (V4 SM120 crash, Apr 2026), CUTLASS #3096 (SM120 TMA WS grouped GEMM failure), and FlashInfer requiring 12 patches to compile on SM120. No ETA as of 2026-05-26.
- Practical recommendations:
- If you already own Pro 5000 cards: SGLang + disable DeepGEMM + Triton backend can pass the mHC layer but likely still fails on attention; BF16/FP16 with TP=4 is tight on 72GB per card; Pro 6000 (96GB) has the same SM120 kernel gap with more VRAM headroom; waiting is the cleanest option with unknown timeline.
- If you are choosing hardware for V4 production: H100/H200 (SM90) or B200 (SM100) have mature vLLM and SGLang support. For local development, 96GB Pro 6000 is more useful than 72GB Pro 5000 because many models are borderline on 72GB. SM120 cards are powerful but best suited to models without custom kernels (Llama 3, standard Qwen3) or to users willing to be early adopters.
- Root cause and historical pattern. Consumer Blackwell has been deprioritized by the inference stack because DeepGEMM-style kernels are written for datacenter SM90/SM100 first. A similar gap occurred when SM89 (RTX 4090) launched. Community patches typically arrive in 3–6 months, but day-one users must accept fallback performance loss or switch to SM90/SM100 hardware.
- DeepGEMM Issue #317: DeepSeek-V4 on SM120 — Unsupported architecture
- DeepGEMM Issue #236: Feature Request: Support sm_120 (5090 and Blackwell 6000 Pro)
- ktransformers Issue #2001: RTX PRO 4000 Blackwell all-zero tokens
- CUTLASS Issue #3096: SM120 TMA WS grouped GEMM failure
- FlashInfer Issue #2577: NVFP4 mm_fp4 GEMM broken on SM120