[论文] Knowing When to Yield: Grounded Arbitration of User Corrections in Tex...
研究领域: ML 作者: Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin 发布时间: 2026-10-05 arXiv: 2610.00282
论文概要
研究领域: ML 作者: Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin 发布时间: 2026-10-05 arXiv: 2610.00282
中文摘要
当人的纠正可能出错,具身智能体该如何回应?我们将有根据的纠正仲裁形式化为接受、拒绝、检查世界、询问说话者间的选择。GAVA 以有界观测证据、合法探针和一步期望损失规则实现该接口。纯文本 ALFWorld 中,162 个检查点产生 972 对真假干预。完整局部检查使 GAVA 与「总是验证」均达 100% 纠正准确率,确立的是证据契约而非比较优势。同回合执行中 GAVA 较「总是验证」降交互成本,完美说话者下与成本阈值持平。仅训练用探索性物体位置先验,在 340 个未见场景上相对均匀 GAVA 将交互成本与声明联合成本分别降 0.490 和 0.420。冻结策略后,收益在 77 个不重叠已见检查点(308 场景)复现:0.595 和 0.517,95% 检查点 bootstrap 区间均排除零。联合成本亦优于同先验固定策略,校准无 VOI 比较仍不确定。语义 GAVA 每队列各犯四个事实错误,对应 98.8% 和 98.7% 准确率,所有方法完成所有任务。结果支持声明成本下用语义先验的选择性信息收集,但不确立环境信息价值对澄清询问的普遍优势。本研究用归一化主张、完整符号观测和受控说话者;未评估人类参与者、视觉输入或物理机器人。
原文摘要
How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, 162 checkpoints produce 972 paired true and false interventions. Complete local inspections give GAVA and always verify 100 percent correction accuracy, establishing the evidence contract rather than a comparative advantage. In same-episode execution, GAVA reduces interaction cost against always verify but ties a cost threshold under a perfect speaker. An exploratory training-only object-location prior lowers interaction and declared joint cost ...
*自动采集于 2026-10-05*
#论文 #arXiv #ML #小凯