[论文] InterEvolve: Test-Time Evolution of Reward Programs for Humanoid Loco-...
研究领域: CV 作者: Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui 发布时间: 2026-10-01 arXiv: 2610.02196
论文概要
研究领域: CV 作者: Zhuo Lin, Sirui Xu, Liuyu Bian, Yu-Xiong Wang, Liang-Yan Gui 发布时间: 2026-10-01 arXiv: 2610.02196
中文摘要
我们研究人形机器人运动-操作的测试时进化:通过重新利用控制器现有技能来完成它从未训练过的任务,从自身尝试中改进并保留所学,全程无需重新训练。核心洞察是:能力广泛的控制器已握有新任务所需的大部分胜任力,而这种胜任力可通过规划与控制之间的接口释放出来——该接口表达力足以刻画富含接触、多阶段的交互,又可执行、可度量,使执行反馈能从经验中引导规划。InterEvolve 用两个组件实现这一接口。其一,物体感知的前向-反向(FB)行为基础模型:其物体残差作用在冻结的身体先验上,可在测试时将关于身体或物体的新奖励转化为运动-操作行为。其二,将任务指定为奖励程序——带完成条件和可调常数的分阶段奖励;LLM 智能体结合执行反馈与经验证的技能程序库在上下文中修订程序结构,数值优化器调节常数。每个候选都在并行仿真中得到验证,程序由此随迭代探索诱导、复用与组合控制器现有运动胜任力的新方式。实验表明:人工设计的奖励让 FB 模型的大部分运动-操作能力沉睡未用,而 InterEvolve 进化出的程序将其释放,有时甚至通过全新策略。它还在仿真中处理多样任务、复杂场景与长时程组合,进化出的技能可基于第一人称机载感知在物理 Unitree G1 上自主运行。
原文摘要
We study test-time evolution for humanoid loco-manipulation: solving tasks that a controller was never trained for by repurposing its existing skills, improving from its own attempts, and retaining what it learns, without retraining. Our key insight is that a broad controller already holds much of the competence a new task needs, and that this competence becomes accessible through an interface between planning and control that is expressive enough to specify contact-rich, multi-stage interactions, yet executable and measurable enough that execution feedback can guide planning from experience. InterEvolve realizes this interface with two components. First, we develop an object-aware forward-backward (FB) behavioral foundation model, whose object residuals on a frozen body prior turn a new r...
*自动采集于 2026-10-04*
#论文 #arXiv #CV #小凯