LLM inference makes KV cache a VRAM nightmare: with long sequences it can reach tens of gigabytes—several times larger than the model weights themselves. Quantizing each value from 16 bits down to 2 bits is the direct fix, but until now nobody has managed to keep 2-bit KV cache quantization accurate.
Simple rotations (e.g., Hadamard transforms) reduce outliers, but at 2-bit precision accuracy still collapses to nearly zero. Zhou, Zhuang, Li, Chen, Song, Athiwaratkun, and Wu found the reason: the rotation is not aligned with the covariance structure consumed by attention—the quantized KV distribution mismatches what attention actually needs to see.
OSCAR's approach:
- Offline, estimate the covariance structure each attention head actually consumes at inference time.
- Use those statistics to derive fixed rotation matrices and clipping thresholds.
- No online computation is required: covariances are computed from sampled calibration data before deployment, and the rotation matrices are frozen.
- A custom INT2 attention CUDA kernel is fully compatible with paged KV cache and fused kernel pipelines, embedding directly into SGLang and vLLM.
- On Qwen3-4B and Qwen3-8B, OSCAR's 2-bit accuracy gap to BF16 is only 1.4–3.8 points, while naive-rotation INT2 collapses to near-zero accuracy.
- Qwen3-32B and GLM-4.7 (358B parameters) likewise stay on par with BF16.
- Long context (128K RULER-NIAH) remains robust.
- System level: ~8x KV cache memory reduction, 7x throughput gain at large batch sizes, and 3x single-sequence decoding speedup.
- How much calibration data does the offline covariance estimation need—how many samples are enough?
- How do covariance structures differ across task types—if the calibration data's distribution differs from the inference-time task distribution, does the rotation matrix remain valid?
- MFU numbers for the INT2 kernel on the 358B model are not disclosed.
Results
Open questions
1. Zhou, Z., Zhuang, D., Li, J., et al. (2026). *OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization*. arXiv:2605.17757 [cs.LG]. 2. Dao, T., et al. (2022). *FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness*. NeurIPS. 3. Ashkboos, S., et al. (2024). *QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks*. ICML.