[论文] Harness Learning Enables Generalizable Test-Time Adaptation

研究领域: NLP 作者: Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar, Daman Arora, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette 发布时间: 2026-09…

论文概要

研究领域: NLP 作者: Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar, Daman Arora, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette 发布时间: 2026-09-28 arXiv: 2609.35738

中文摘要

语言模型智能体由其模型和外部框架(harness)联合定义——harness 是组织模型调用、工具使用和信息流的可执行程序。因为不同任务需要不同的组织方式,harness 需要利用任务反馈进行适配。我们引入 harness 学习,训练一个提议模型利用执行反馈来修正求解器的 harness。我们将此过程形式化为可执行程序上的元学习,harness 修正扮演梯度适配中权重更新的角色。我们使用强化学习训练提议模型,以修正后 harness 的任务性能作为奖励。测试时,提议模型利用来自新任务上连续执行的反馈来精炼 harness,而不进行任何参数空间更新。在推理和多跳问答上的实验表明,harness 学习改善了修正质量,且测试时适配能力可迁移到未见任务。在单项修正上训练的策略可以在多轮中持续改进 harness。这些发现指向一条通往持续学习智能体的路径——将积累的经验转化为可泛化的改进。

原文摘要

A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, wit...


*自动采集于 2026-09-30*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens