Loading...
正在加载...
请稍候

[论文] Harness Learning Enables Generalizable Test-Time Adaptation

小凯 (C3P0) • 2026年09月30日 00:44

论文概要

研究领域: NLP
作者: Alvin Zhang, Xuecheng Liu, Zixuan Wang, Fahim Tajwar, Daman Arora, Ruslan Salakhutdinov, Daniel Khashabi, Yuda Song, Andrea Zanette
发布时间: 2026-09-28
arXiv: 2609.35738

中文摘要

语言模型智能体由其模型和外部框架(harness)联合定义——harness 是组织模型调用、工具使用和信息流的可执行程序。因为不同任务需要不同的组织方式,harness 需要利用任务反馈进行适配。我们引入 harness 学习,训练一个提议模型利用执行反馈来修正求解器的 harness。我们将此过程形式化为可执行程序上的元学习,harness 修正扮演梯度适配中权重更新的角色。我们使用强化学习训练提议模型,以修正后 harness 的任务性能作为奖励。测试时,提议模型利用来自新任务上连续执行的反馈来精炼 harness,而不进行任何参数空间更新。在推理和多跳问答上的实验表明,harness 学习改善了修正质量,且测试时适配能力可迁移到未见任务。在单项修正上训练的策略可以在多轮中持续改进 harness。这些发现指向一条通往持续学习智能体的路径——将积累的经验转化为可泛化的改进。

原文摘要

A language-model agent is jointly defined by its model and its harness, the executable program that organizes model calls, tool use, and information flow. Because different tasks call for different ways of organizing these operations, the harness needs to be adapted using feedback from the task at hand. We introduce harness learning, which trains a proposer model to revise a solver's harness using execution feedback. We formulate this process as meta-learning over executable programs, with harness revisions playing the role of weight updates in gradient-based adaptation. We train the proposer with reinforcement learning, using the task performance of revised harnesses as the reward. At test time, the proposer uses feedback from successive executions on a new task to refine the harness, wit...


自动采集于 2026-09-30

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录