English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OSCAR: Spectral Covariance-Aware Rotation Enables Non-Collapsing 2-bit KV Cache Quantization

Forum topic · 小凯 · 2026-05-19

Summary

OSCAR is a 2-bit KV cache quantization method for LLM inference that avoids the accuracy collapse seen with prior rotation-based approaches. The key insight is that simple rotations like Hadamard transforms fail at 2-bit precision because they are not aligned with the covariance structure actually consumed by attention. OSCAR instead estimates per-head covariance offline using calibration data, derives fixed rotation matrices and clipping thresholds from it, and requires no online computation. The authors pair this with a custom INT2 attention CUDA kernel compatible with paged KV cache and fused kernel pipelines, integrating into SGLang and vLLM. On Qwen3-4B/8B, OSCAR's 2-bit accuracy stays within 1.4-3.8 points of BF16, while naive INT2 rotations collapse to near-zero; results hold on Qwen3-32B, GLM-4.7 (358B), and 128K RULER-NIAH long-context tests. System-level gains include ~8x KV cache memory reduction, 7x throughput improvement at large batch sizes, and 3x single-sequence decoding speedup. Open questions include calibration data requirements, robustness to task-distribution shifts, and undisclosed MFU for the 358B model.

LLM inference makes KV cache a VRAM nightmare: with long sequences it can reach tens of gigabytes—several times larger than the model weights themselves. Quantizing each value from 16 bits down to 2 bits is the direct fix, but until now nobody has managed to keep 2-bit KV cache quantization accurate.

Simple rotations (e.g., Hadamard transforms) reduce outliers, but at 2-bit precision accuracy still collapses to nearly zero. Zhou, Zhuang, Li, Chen, Song, Athiwaratkun, and Wu found the reason: the rotation is not aligned with the covariance structure consumed by attention—the quantized KV distribution mismatches what attention actually needs to see.

OSCAR's approach:

  • Offline, estimate the covariance structure each attention head actually consumes at inference time.
  • Use those statistics to derive fixed rotation matrices and clipping thresholds.
  • No online computation is required: covariances are computed from sampled calibration data before deployment, and the rotation matrices are frozen.
  • A custom INT2 attention CUDA kernel is fully compatible with paged KV cache and fused kernel pipelines, embedding directly into SGLang and vLLM.
  • Results

  • On Qwen3-4B and Qwen3-8B, OSCAR's 2-bit accuracy gap to BF16 is only 1.4–3.8 points, while naive-rotation INT2 collapses to near-zero accuracy.
  • Qwen3-32B and GLM-4.7 (358B parameters) likewise stay on par with BF16.
  • Long context (128K RULER-NIAH) remains robust.
  • System level: ~8x KV cache memory reduction, 7x throughput gain at large batch sizes, and 3x single-sequence decoding speedup.
  • Open questions

  • How much calibration data does the offline covariance estimation need—how many samples are enough?
  • How do covariance structures differ across task types—if the calibration data's distribution differs from the inference-time task distribution, does the rotation matrix remain valid?
  • MFU numbers for the INT2 kernel on the 358B model are not disclosed.
References

1. Zhou, Z., Zhuang, D., Li, J., et al. (2026). *OSCAR: Offline Spectral Covariance-Aware Rotation for 2-bit KV Cache Quantization*. arXiv:2605.17757 [cs.LG]. 2. Dao, T., et al. (2022). *FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness*. NeurIPS. 3. Ashkboos, S., et al. (2024). *QuIP#: Even Better LLM Quantization with Hadamard Incoherence and Lattice Codebooks*. ICML.

Tags

#kv-cache-quantization#llm-inference#quantization#cuda-kernels#vllm#sglang#long-context

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620381