[论文] Knowing When to Yield: Grounded Arbitration of User Corrections in Tex...

研究领域: ML 作者: Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin 发布时间: 2026-10-05 arXiv: 2610.00282

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Yezhou Cheng, Runjia Du, Zeming Liu, Hang Lyu, Zehua Yang, Bojun Lin 发布时间: 2026-10-05 arXiv: 2610.00282

中文摘要

当人的纠正可能出错,具身智能体该如何回应?我们将有根据的纠正仲裁形式化为接受、拒绝、检查世界、询问说话者间的选择。GAVA 以有界观测证据、合法探针和一步期望损失规则实现该接口。纯文本 ALFWorld 中,162 个检查点产生 972 对真假干预。完整局部检查使 GAVA 与「总是验证」均达 100% 纠正准确率,确立的是证据契约而非比较优势。同回合执行中 GAVA 较「总是验证」降交互成本,完美说话者下与成本阈值持平。仅训练用探索性物体位置先验,在 340 个未见场景上相对均匀 GAVA 将交互成本与声明联合成本分别降 0.490 和 0.420。冻结策略后,收益在 77 个不重叠已见检查点(308 场景)复现:0.595 和 0.517,95% 检查点 bootstrap 区间均排除零。联合成本亦优于同先验固定策略,校准无 VOI 比较仍不确定。语义 GAVA 每队列各犯四个事实错误,对应 98.8% 和 98.7% 准确率,所有方法完成所有任务。结果支持声明成本下用语义先验的选择性信息收集,但不确立环境信息价值对澄清询问的普遍优势。本研究用归一化主张、完整符号观测和受控说话者;未评估人类参与者、视觉输入或物理机器人。

原文摘要

How should an embodied agent respond when a person's correction may be wrong? We formulate grounded correction arbitration as a choice among accepting, rejecting, inspecting the world, and asking the speaker. GAVA implements this interface with observation-bounded evidence, legal probes, and a one-step expected-loss rule. In text-only ALFWorld, 162 checkpoints produce 972 paired true and false interventions. Complete local inspections give GAVA and always verify 100 percent correction accuracy, establishing the evidence contract rather than a comparative advantage. In same-episode execution, GAVA reduces interaction cost against always verify but ties a cost threshold under a perfect speaker. An exploratory training-only object-location prior lowers interaction and declared joint cost ...


*自动采集于 2026-10-05*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens