[论文] VeriFine: Scaling Verification for Self-Improvement in Embodied Reason...

研究领域: ML 作者: Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang, Jiaqi Ma, Boris Ivanovic, Marco Pavo…

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang, Jiaqi Ma, Boris Ivanovic, Marco Pavone, Wenhao Ding 发布时间: 2026-10-06 arXiv: 2610.08761

中文摘要

自我改进的策略不断暴露新的失败模式,改变其评判者必须能够验证的内容。然而,当前的固定评判者同时限制了优化反馈和有用训练示例的发现,限制了进一步的自我改进。这一挑战在具身推理中更加突出,其中可靠的评估必须考虑空间定位、因果推理和安全感知决策。我们引入VeriFine,这是一个智能体框架,通过策略、训练课程和评判者的共同演化来扩展验证。策略改进循环使用评分标准评判者来诊断反复出现的失败,构建自适应课程,并优化策略。当进展停滞且验证成为瓶颈时,评判者改进循环选择性地查询关于信息性失败案例的人类指导,并通过交互校准来细化评判者——人类和智能体在其中解决分歧并收敛到物理推理的目标评分标准。修订后的评判者然后指导下一阶段的数据选择和策略优化。在驾驶和机器人导航任务上的实验表明,在强化和supervised fine-tuning中,策略和评判者能力都实现了持续自我改进。

原文摘要

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improv...


*自动采集于 2026-10-08*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens