[论文] Rephrase Before You Act: Characterizing and Mitigating Language Sensit...

研究领域: NLP 作者: Mikey Watts, Yuchen Cui 发布时间: 2026-10-07 arXiv: 2610.10526

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: NLP 作者: Mikey Watts, Yuchen Cui 发布时间: 2026-10-07 arXiv: 2610.10526

中文摘要

视觉-语言-动作模型(VLA)对指令措辞极为敏感,且并未继承其底层视觉-语言模型的语言鲁棒性。一个词的修改就能使成功率波动数十个百分点:π0.5对"switch on the stove"的开 stove 成功率为100%,而对"switch on the hot plate"仅为2%;一个经过重述增强微调的π0检查点仍表现出高达61个百分点的波动。我们通过统计检验的单编辑摆动和神谕短语搜索来刻画这种敏感性——后者表明仅靠措辞优化就能弥合分布内与分布外任务之间21个百分点的差距。我们在不修改策略的情况下降低这种敏感性:由于敏感性是系统性的,可以表示为显式规则——我们对少量训练任务的多种措辞进行评分,让大语言模型将证据蒸馏为10到20条重述规则,部署时将每条传入指令按这些规则重写一次。这些规则在12个留出任务上将冻结π0的相对性能提升16%到27%(涵盖对抗性、VLM生成和人类生成的措辞),增益集中在分布外任务上。该方法在π0.5和LIBERO上同样有效,将微调内成功率从93.6%提升至97.8%。无需重训练、无需逐步验证,可零样本应用于未见任务和指令。

原文摘要

Vision-language-action models (VLAs) are strikingly sensitive to instruction phrasing and do not inherit the language robustness of the vision-language models they are built on. A one-word edit can move success by tens of points: \(π_{0.5}\) turns on a LIBERO stove 100% of the time for "switch on the stove" and 2% for "switch on the hot plate", and a \(π_0\) checkpoint finetuned with rephrase augmentation still shows swings of up to 61 points. We characterize this sensitivity with statistically tested single-edit swings and an oracle phrase search, which shows that phrasing alone nearly closes the 21-point gap between in-distribution and out-of-distribution tasks. We then reduce it without modifying the policy. Because the sensitivity is systematic, it can be expressed as explicit rules: we sc...


*自动采集于 2026-10-09*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens