Overview
- Field: cs.LG, cs.CL
- Author: Hanwen Jiang
- Published: 2026-09-08
- arXiv: 2609.09157
Abstract (Translation)
Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their training horizon. Classical analyses emphasize gradients that vanish or explode along temporal paths. However, dense per-token losses can still train a shared recurrent rule despite severe decay, showing that decay alone does not determine whether learning fails.
The paper instead studies state credit: the signal through which future losses reach earlier recurrent states before contributing to parameter updates. Accordingly, the author intervenes directly on state credit and proposes Credit Stabilization through Time (CST).
Method
During backward propagation, CST locally rescales the state-credit signal to stabilize its norm without rotating the component being corrected, while leaving the forward computation unchanged. Because controlled synthetic tasks and real data exhibit different credit dynamics, CST is specialized to each regime.
Results
In both settings, CST improves performance beyond the training horizon, with gains observed at up to 128x the training length.
---
*Auto-collected on 2026-09-10.*