[论文] Same Trajectory, Contradictory Rewards (ROBORMBENCH): Paraphrase Fragi...

研究领域: NLP 作者: Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No 发布时间: 2026-09-04 arXiv: 2609.05401

论文概要

研究领域: NLP 作者: Wonje Jeung, Sangyeon Yoon, Hyesoo Hong, Yoonjun Cho, Dongjae Jeon, Bumjun Kim, Jean Oh, Youngjae Yu, Albert No 发布时间: 2026-09-04 arXiv: 2609.05401

中文摘要

翻译缺失

原文摘要

Vision-language models are increasingly used as reward functions for robotic learning, but this role requires paraphrase invariance: the same trajectory should receive the same reward under semantically equivalent goal descriptions. We show that current VLM reward models often violate this property. Paraphrasing the instruction alone can substantially change predicted progress scores, and can even flip identical robot behavior between failure and success. To measure this failure mode, we introduce ROBORMBENCH, a benchmark with 2,390 real-robot trajectories, ground-truth progress labels, and 21,673 verified paraphrases spanning lexical, syntactic, and action-goal rewrites. Across proprietary and open-source VLMs, paraphrase-induced instability is widespread and severe, grows under more dive...


*自动采集于 2026-09-08*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(1)

Q

这篇的 arXiv 号难得是对的:2609.05401,9 月 4 号挂出,延世大学为主力,CMU 的 Jean Oh 也在作者列表里。帖子里写研究领域是 NLP,其实归档在 cs.RO,机器人。

论文里有个例子我看了三遍。同一段机器人操作录像,目标写"Pick radish, then place in pink bowl",模型给 1 分;改成"After picking radish, place in pink bowl",同一段录像给 5 分。萝卜还是那根萝卜。碗还是那个粉碗。同一个模型,分数从失败跳到满分。他们攒了 2,390 条真机轨迹、21,673 条验证过的改写,专门称这种事:改写引起的分数波动标准差 0.52,同一个提示词反复问的噪声只有 0.09。改写的杀伤力是噪声的六倍。

翻得最狠的是动作-目标类改写。Gemini2.5-flash-lite 和 Llama4-scout 在这类改写下,过半轨迹的成功判定会被直接翻转。堆参数救不了,让模型显式推理也救不了,论文原话是"not reliably reduced by scale or explicit reasoning"。最稳的通用模型是 Claude Sonnet 4.6,专训的 RoboReward 4B/8B 比通用 VLM 稳约十倍。

真正管用的药在训练里:方差抑制训练把 Qwen3-VL-4B 的不稳定分数从 0.076 压到 0.042,平均误差从 1.046 砍到 0.239。提示词工程全数阵亡,训练目标一改就见效。

一个瑕疵:论文宣称基准按 CC BY 4.0 发布。我翻了全文再搜全网,仓库链接不存在。想复现的再等等吧。

暂无表态

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens