[论文] Learning Length-Extrapolatable Recurrent Models

论文概要 研究领域: cs.LG, cs.CL 作者: Hanwen Jiang 发布时间: 2026-09-08 arXiv: 2609.09157

论文概要

研究领域: cs.LG, cs.CL 作者: Hanwen Jiang 发布时间: 2026-09-08 arXiv: 2609.09157

中文摘要

循环模型为长上下文建模提供了天然路径,然而通过时间反向传播(BPTT)训练的模型通常在训练范围之外失效。经典分析强调梯度沿时间路径消失或爆炸。但即使在严重衰减的情况下,密集的逐token损失仍能训练共享的循环规则,这表明衰减本身并不能决定学习是否失败。我们转而研究状态信用:未来损失到达较早循环状态并贡献于参数更新的信号。据此,我们直接干预状态信用,提出时间信用稳定化(CST)。在反向传播期间,CST局部重新缩放状态信用信号以稳定其范数,同时不改变被修正分量的方向,且保持前向计算不变。由于受控合成任务和真实数据表现出不同的信用动态,我们为每种机制专门设计了CST。在两种设置中,CST均提升了训练范围之外的性能,在长达训练长度128倍的范围内观察到增益。

原文摘要

Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails. We instead study state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, we intervene directly on state credit and propose Credit Stabilization through Time (CST). During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, we specialize CST to each regime. In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.


*自动采集于 2026-09-10*

#论文 #arXiv #AI #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens