Loading...
正在加载...
请稍候

[论文] Rephrase Before You Act: Characterizing and Mitigating Language Sensit...

小凯 (C3P0) • 2026年10月09日 00:43

论文概要

研究领域: NLP
作者: Mikey Watts, Yuchen Cui
发布时间: 2026-10-07
arXiv: 2610.10526

中文摘要

视觉-语言-动作模型(VLA)对指令措辞极为敏感,且并未继承其底层视觉-语言模型的语言鲁棒性。一个词的修改就能使成功率波动数十个百分点:π0.5对"switch on the stove"的开 stove 成功率为100%,而对"switch on the hot plate"仅为2%;一个经过重述增强微调的π0检查点仍表现出高达61个百分点的波动。我们通过统计检验的单编辑摆动和神谕短语搜索来刻画这种敏感性——后者表明仅靠措辞优化就能弥合分布内与分布外任务之间21个百分点的差距。我们在不修改策略的情况下降低这种敏感性:由于敏感性是系统性的,可以表示为显式规则——我们对少量训练任务的多种措辞进行评分,让大语言模型将证据蒸馏为10到20条重述规则,部署时将每条传入指令按这些规则重写一次。这些规则在12个留出任务上将冻结π0的相对性能提升16%到27%(涵盖对抗性、VLM生成和人类生成的措辞),增益集中在分布外任务上。该方法在π0.5和LIBERO上同样有效,将微调内成功率从93.6%提升至97.8%。无需重训练、无需逐步验证,可零样本应用于未见任务和指令。

原文摘要

Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: \(π_{0.5}\) turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a \(π_0\) checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we sc...


自动采集于 2026-10-09

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录