[论文] Same evidence, different judgments: Evidence noncommutative in vision/...

研究领域: ML 作者: Zhuoyun Li, Boxuan Wang, Xiaowei Huang, Yi Dong 发布时间: 2026-09-25 arXiv: 2609.26986

论文概要

研究领域: ML 作者: Zhuoyun Li, Boxuan Wang, Xiaowei Huang, Yi Dong 发布时间: 2026-09-25 arXiv: 2609.26986

中文摘要

对多模态大语言模型而言,当图像或语音与随附文本冲突时,测得的文本依赖度可能将模态偏好与证据位置纠缠在一起。早期文本偏置研究常用固定证据顺序,或让任务指令随证据移动,导致顺序的贡献不明确。本文使用配对比较:保持指令与证据内容不变,仅交换两个来源的位置,量化这种潜在影响。在视觉与语音模型中,将图像或录音置于冲突文本之后会持续把答案推向其内容。我们还重审先前研究,分析其实验设定为何可能得出误导结论。这些发现揭示了跨模态证据的不可交换性:同样证据在顺序变化时可导致不同判断;将感知证据放在更靠后的位置会增强模型对其内容的依赖。

原文摘要

For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unclear. In this paper, we use a paired comparison that keeps the instructions and evidence content fixed and swaps only the positions of the two sources to quantify this potential influence. Across vision and speech models, placing an image or recording after conflicting text consistently shifts answers toward its content. We also revisit previous studies and analyze why their experimental settings can lead to misleading conclusions. These findings reveal cross-modal evidence no...


*自动采集于 2026-09-25*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens