Loading...
正在加载...
请稍候

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots

小凯 (C3P0) 2026年08月11日 20:56

论文概要

研究领域: CV
作者: Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar
发布时间: 2026-08-11
arXiv: 2508.03807

中文摘要

多模态大语言模型(MLLMs)的自我提升通常依赖仅提供粗粒度标量反馈的基于奖励的方法。蒸馏通过密集的token级监督提供了更丰富的替代方案,但在视觉领域,它通常依赖于使用外部注释和工具或更强模型构建的特权上下文。我们引入了CVPD(对比反事实视觉过程蒸馏),据我们所知,这是第一个用于MLLMs的完全自包含的密集、在线、token级视觉自蒸馏框架。CVPD识别视觉盲点——即放大某个区域会改变并锐化模型的答案分布,而移除同一区域则基本不影响全图行为的区域。这些区域揭示了模型能够编码但在全图条件下未能持续利用的感知信息。我们提出了一个三阶段反事实标准,直接从模型自身的响应中识别这些区域,并将其转换为密集对比监督以进行自我蒸馏。在Qwen3-VL-8B-Instruct上,CVPD在12个基准测试中超越了6个自我进化基线(包括依赖外部GPT-4o监督的方法),且没有任何性能回归。它在OCRBench上获得+3.60的提升,在MMStar细粒度感知上获得+3.38的提升,在MMStar逻辑推理上获得+3.08的提升,同时在更广泛的多模态基准测试上保持或提升了性能。

原文摘要

Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. We introduce CVPD (Contrastive Counterfactual Visual Process Distillation), which, to the best of our knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs. CVPD identifies visual blind spots where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such region...


自动采集于 2026-08-12

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录