[论文] Improving Test-Time Scaling with Adaptive Looped Transformers
研究领域: NLP 作者: Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang 发布时间: 2026-09-28 arXiv: 2609.35748
论文概要
研究领域: NLP 作者: Yichen You, Tianyu Fu, Aosong Feng, Xingtai Lv, Xuefei Ning, Ning Ding, Yu Wang 发布时间: 2026-09-28 arXiv: 2609.35748
中文摘要
循环 Transformer 通过复用层进行潜在计算,展示了有前景的参数效率。先前研究在匹配参数或每 token FLOPs 下比较循环和非循环模型。然而,据我们所知,随着输出变长,循环是否改善测试时缩放仍有待探索。通过对循环 Transformer 进行后训练,我们研究了准确率-计算斜率,即测试时解码 FLOPs 每翻倍的准确率增益。我们发现现有循环 Transformer 通常比非循环基线产生更陡的斜率,但在匹配计算下表现更差。固定深度循环在每个 token 上花费额外迭代,但我们的分析表明许多 token 并不从额外迭代中受益。因此我们提出 TaH2,使模型能将额外迭代集中在受益于循环的 token 上。它通过前瞻深度监督联合后训练骨干网络和迭代决策者,使用在线标签指示进一步迭代是否改善预测。TaH2 改善了测试时缩放的效率和可达准确率。在 AIME 基准上,TaH2 将准确率-计算斜率比非循环基线提高 53%(2.74 vs. 1.79),在匹配测试时计算下超过基线峰值准确率约 3.4 分。随着最大迭代深度增加,现有循环模型基本趋于平稳,而 TaH2 相对于基线的增益从深度 2 的 +2.8 分持续增长到深度 8 的 +3.9 分。
原文摘要
Looped transformers have demonstrated promising parameter efficiency by reusing layers for latent computation. Prior studies compare looped and non-looped models at matched parameters or per-token FLOPs. However, to the best of our knowledge, whether looping improves test-time scaling as outputs grow longer remains underexplored. Through post-training looped transformers, we study the accuracy-compute slope, measured as the accuracy gain per doubling of test-time decoding FLOPs. We find that existing looped transformers often yield steeper slopes than their non-looped baseline, yet underperform it at matched compute. While fixed-depth looping spends extra iterations on every token, our analysis shows that many tokens do not benefit from extra iterations. We therefore propose TaH2, which en...
*自动采集于 2026-09-30*
#论文 #arXiv #NLP #小凯