论文概要
研究领域: ML
作者: Mingxuan Wang, Bo Wang, Fei Luo, Guorun Yao, Chao Ning, Yinglong Guo, Hongyue Chen, Yanbiao Ma, Jungong Han
发布时间: 2026-09-25
arXiv: 2609.27276
中文摘要
长程语言模型智能体会积累推理轨迹、工具交互与观察,其相关性随当前决策变化。现有压缩策略通常独立地对历史单元打分,但删除多个单元的安全性一般并非由各自单独分数决定:冗余证据、累积的小效应,以及删除后残留的信息都很关键。我们提出 DRSR(直接关系集合风险剪枝),把智能体历史压缩表述为对删除集合的风险约束选择。离线阶段,DRSR 通过联合删除协议合法的历史块、测量同一已记录下一输出的教师强制似然变化,构造精确反事实监督。轻量打分器随后从在线可见的候选历史与当前动作前状态之间的关系,连同删除-保留与成对集合结构,预测集合层面伤害。部署时,DRSR 用轻量打分器评估一小组建构合法的删除候选,在新近性、协议、预算与习得风险约束下移除最大可行集合,若无足够安全的集合则弃权。WorkBuddyBench Full260 上,DRSR 将平均奖励从 0.699 提升至 0.802,模型总 token 减少 20.820%。固定 Eval40 比较上,它以每任务 1.211M token 获得 0.794 奖励,比未压缩智能体少用 35.850% token。机制分析与消融表明:决策条件关系、保留上下文信息、成对交互与弃权机制各自都对可靠剪枝有贡献。
原文摘要
Long-horizon language-model agents accumulate reasoning traces, tool exchanges, and observations whose relevance changes with the current decision. Existing compression strategies often score historical units independently, but the safety of deleting several units is generally not determined by their singleton scores: redundant evidence, accumulated small effects, and the information that remains after deletion all matter. We introduce Direct Relational Set-Risk Pruning (DRSR), which formulates agent-history compression as risk-constrained selection over deletion sets. Offline, DRSR constructs exact counterfactual supervision by jointly deleting protocol-valid history Blocks and measuring the change in teacher-forced likelihood of the same recorded next output. A lightweight scorer then pr...
自动采集于 2026-09-25
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。