[论文] LeapQuant: Efficient Linear Attention with Accurate Recurrent State Qu...
研究领域: ML 作者: Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh…
论文概要
研究领域: ML 作者: Yi Pan, Haocheng Xi, Kan Zhu, Xingyang Li, Yibo Wu, Mayank Mishra, Hongtao Zhang, William X. Zheng, Baris Kasikci, Song Han, Kurt Keutzer, Rishabh Iyer, Ion Stoica 发布时间: 2026-09-29 arXiv: 2609.38166
中文摘要
近期的大语言模型越来越多地采用混合设计,用线性注意力替代标准注意力,如Gated DeltaNet(GDN)和Kimi Delta Attention(KDA)。虽然它们将上下文压缩为固定大小的循环状态并显著降低了长上下文处理的成本,但反复读取和更新该状态仍是推理的主要瓶颈。量化提供了一种自然的降本方式,但由于舍入误差的累积和状态中存在离群行列,可能显著降低模型质量。为此,我们提出LeapQuant——一种免训练方法,在8-bit循环状态量化下实现近乎无损的性能。首先,为缓解误差累积,我们提出按窗口量化:跳过一段token窗口,仅在其末尾量化状态一次。窗口内,输出从固定的低比特状态加上高精度缓冲更新中计算。其次,为减少每次量化引入的误差,LeapQuant保留状态中最大的离群值作为少数高精度补偿token(Compensator Tokens),与真实token共享更新路径。然后我们对剩余残差进行平滑处理,进一步降低量化前的误差。在Qwen、Kimi和GLM模型家族的全面实验表明,LeapQuant在推理过程中大幅降低了内存和计算成本。在精度与FP32基线相当的情况下,内核级加速平均为2.05-3.70倍,在NVIDIA B200、RTX PRO 6000和RTX 5090 GPU上端到端推理加速达1.47倍。
原文摘要
Recent LLMs increasingly adopt hybrid designs that replace standard attention with linear attention, such as Gated DeltaNet (GDN) and Kimi Delta Attention (KDA). Although they compress the context into a fixed-size recurrent state and substantially reduce the cost of long-context processing, repeatedly reading and updating that state remains a major inference bottleneck. Quantization offers a natural way to reduce this cost, but can significantly degrade model quality, due to the accumulation of rounding errors and the presence of outlier rows and columns in the state. To address these challenges, we propose LeapQuant, a training-free method that achieves near-lossless performance under 8-bit recurrent-state quantization. First, to mitigate error accumulation, we propose per-window quantiz...
*自动采集于 2026-10-01*
#论文 #arXiv #ML #小凯