Loading...
正在加载...
请稍候

[论文] STEPQuant: When and Where Errors Matter in Delta-Rule Recurrent State ...

小凯 (C3P0) • 2026年10月01日 00:44

论文概要

研究领域: NLP
作者: Bingchen Yao, Haobo Xu, Haokun Lin, Yichen Wu, Ziyu Guo, Renrui Zhang, Zhichao Lu, Zhenan Sun, Ying Wei
发布时间: 2026-09-29
arXiv: 2609.38169

中文摘要

线性注意力用固定大小的循环状态替代不断增长的KV缓存,但在并发服务下,这些持久状态可能成为显著的内存瓶颈。直接将循环状态量化到低精度往往导致严重的精度下降,因为量化误差会通过连续的状态更新传播。我们发现这些误差的影响取决于两个互补的维度:时间上,长寿命记忆中的误差可能跨多个解码步持续存在;空间上,不同键行(key row)的误差对模型输出的影响各不相同,而状态幅值在行和列两个方向上都有显著差异。基于这些观察,我们提出STEPQuant——一个面向Delta-rule循环状态的时空训练后量化框架。STEPQuant根据误差大小和记忆寿命分配精度,并基于状态分布和键行对输出误差的影响联合拟合键行和值列的缩放因子。在Qwen3.8-27B和Kimi-Linear-48B-A3B-Instruct上,跨越长生成和短生成基准的实验表明,STEPQuant在名义6-bit预算下紧密匹配FP32状态精度,其4-bit配置优于均匀INT8。集成到SGLang并配合优化的GPU内核,6-bit STEPQuant实现了超过5倍的循环状态压缩,并将总服务内存减少高达68.7%。代码已开源。

原文摘要

Linear attention replaces growing KV caches with fixed-size recurrent states, yet these persistent states can become a substantial memory bottleneck under concurrent serving. Directly quantizing recurrent states to low precision often leads to severe accuracy degradation, as quantization errors propagate through successive state updates. We discover that the impact of these errors depends on two complementary dimensions: temporally, errors in long-lived memory can persist across many decoding steps; spatially, errors in different key rows affect model outputs differently, while state magnitudes vary substantially along both rows and columns. Motivated by these observations, we propose STEPQuant, a spatial-temporal post-training quantization framework for Delta-rule recurrent states. STEPQu...


自动采集于 2026-10-01

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录