Loading...
正在加载...
请稍候

[论文] DiaVLo: Diagnosing Behaviours of Vision-Language Models

小凯 (C3P0) 2026年09月22日 00:46

论文概要

研究领域: NLP
作者: Lorenzo Corti, Jie Yang
发布时间: 2026-09-18
arXiv: 2609.22008

中文摘要

视觉-语言模型(VLM)依赖在其子组件间存储和传递适当的信息。验证 VLM 表现出期望行为、同时避免有害行为,是其可靠部署的核心。然而,识别 VLM 行为的方法仍然稀缺。我们提出 DiaVLo,一个诊断框架,利用人工策划与 VLM 的生成能力来构建期望行为与观测行为的规约,暴露潜在的对齐偏差。除此之外,DiaVLo 还提供因果估计,用于识别引导 VLM 行为的最有影响力的概念。我们在若干开源 VLM 上、于分类与生成两种条件下评估 DiaVLo。实验表明,DiaVLo 产生的行为标签与模型性能相关,并为测得的性能提供背景信息。DiaVLo 揭示了明显对齐和明显偏差的行为,以及 VLM 感知、组织和优先化概念的模式。

原文摘要

Vision-language models (VLMs) rely on storing and transferring appropriate information across their sub-components. Verifying that the VLMs exhibit desired behaviours, while avoiding harmful ones, is central to their reliable deployment. Yet, methods that identify VLM behaviours remain scarce. We present DiaVLo, a diagnostic framework that leverages human curation and VLMs' generation capabilities to construct specifications of desired and observed VLM behaviours, surfacing potential misalignments. Beyond this, DiaVLo also provides causal estimates to identify the most influential concepts steering VLM behaviours. We evaluate DiaVLo on several open-source VLMs under both classification and generation conditions. Our experiments show that DiaVLo produces behaviour labels that correlate with...


自动采集于 2026-09-22

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录