[论文] KV-streams for Efficient Compaction in Agentic Reinforcement Learning

研究领域: ML 作者: Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sarg…

论文概要

研究领域: ML 作者: Emiliano Penaloza, Dane Malenfant, Dheeraj Vattikonda, Roger Creus Castanyer, Siddarth Venkatraman, Abhay Puri, Jonathan Light, Matthew James Sargent, Augustine N. Mavor-Parker, Massimo Caccia, Lucas Caccia, Glen Berseth, Esmeralda S. Whitammer, Alessandro Sordoni, Minseon Kim, Marc-Alexandre Côté, Laurent Charlin, Guillaume Lajoie 发布时间: 2026-09-28 arXiv: 2609.35750

中文摘要

扩展智能体 LLM 的视野长度的瓶颈在于需要将不断增长的上下文轨迹装入 GPU 显存。上下文压缩一直是缓解此问题的最流行机制,能在给定轨迹下保持 GPU 显存恒定。不幸的是,大多数压缩策略依赖多次预填充 LLM 上下文,阻碍了训练吞吐量。为缓解这一瓶颈并实现高效可训练的压缩,我们提出 KV-streams——一种与任何压缩策略兼容的即插即用策略,可显著提高吞吐量且没有证据表明会损害性能。KV-streams 通过在前向传播中流式传输 KV 缓存而非在每次压缩后将其清空来实现可扩展压缩。我们展示了 KV-streams 支持三种不同的压缩策略,实现了 2.6 到 5 倍的训练加速。除效率外,我们发现流式 KV 缓存可以作为循环状态,携带早已从上下文中消失的信息。具体而言,在受控环境中我们证明,与先前工作不同,仅靠 RL 就足以让这种行为涌现。总体而言,我们表明 KV-streams 是一种高效轻量的即插即用附加组件,适用于任何后训练流水线。

原文摘要

Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular mechanism to alleviate this issue, keeping GPU memory constant for a given trace. Unfortunately, most compaction strategies rely on prefilling the LLM context many times over, hindering training throughput. To alleviate this bottleneck and enable efficient trainable compaction, we propose KV-streams, a plug-and-play strategy compatible with any compaction strategy that substantially increases throughput while showing no evidence of hindering performance. KV-streams enable scalable compaction by streaming the KV cache forward rather than flushing it after each compaction. We show that KV-streams enable three different compaction strate...


*自动采集于 2026-09-30*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens