[论文] Pretraining Latent Information Feedback Transformers with Teacher Supe...
研究领域: NLP 作者: Dor Tirosh, Ido Amos, Mor Geva 发布时间: 2026-09-29 arXiv: 2609.38149
论文概要
研究领域: NLP 作者: Dor Tirosh, Ido Amos, Mor Geva 发布时间: 2026-09-29 arXiv: 2609.38149
中文摘要
Transformer语言模型是前馈的:深层表示从不反馈到浅层,信息跨生成步向下流动的唯一途径是已解码的token。这一狭窄通道迫使模型重新计算中间结果并丢弃替代延续方案。本文在预训练期间移除了这一瓶颈,引入LIFT(潜在信息反馈Transformer)架构和训练方法,使语言模型能够在生成过程中传播状态。我们通过将循环状态学习转化为教师强制的预测问题来实现这一点:每个输入token与一个信息密集的状态配对,该状态来自现成的预训练语言模型的下一token分布。扩展了少量额外参数的模型随后被训练来同时预测下一个token和下一个状态。由于输入状态是预计算的,预训练在位置上完全并行。推理时,模型自身预测的状态被反馈回来,额外计算开销随模型规模增大而减小。在135M到1B参数的预训练模型上的实验表明,在token匹配预算下,LIFT一致优于标准Transformer和基线方法,在语言建模、下游推理任务和程序性任务上表现出色,同时在计算匹配的比较中与Transformer持平或领先。此外,在一个状态追踪任务上的对照研究表明,即使使用一个在该任务上失败的Transformer的状态进行训练,微小的LIFT也优于在8倍数据上训练的同等规模Transformer。总体而言,我们表明语言模型可以通过可扩展的教师监督在预训练期间学会利用从深到浅的反馈。
原文摘要
Transformer language models (LMs) are feed-forward: deep-layer representations are never fed back to shallower layers, and the only pathway for information to flow downward across generation steps is the decoded token. This narrow channel forces models to recompute intermediate results and to discard alternative continuations. In this work, we remove this bottleneck during pretraining, introducing the LIFT (Latent Information Feedback Transformer) architecture and training method which enable LMs to propagate state across generation. We achieve this by turning recurrent-state learning into a teacher-forced prediction problem: each input token is paired with an information-dense state, derived from the next-token distribution of an off-the-shelf pretrained LM. The model, extended with a sma...
*自动采集于 2026-10-01*
#论文 #arXiv #NLP #小凯