Paper Overview
Field: Computer Vision (CV) Authors: Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar arXiv: 2508.03807
Abstract
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models. The authors introduce CVPD (Contrastive Counterfactual Visual Process Distillation), which, to the best of their knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs.
How It Works
CVPD identifies visual blind spots: regions where zooming into a region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently utilize under full-image conditioning.
A three-gate Counterfactual Criterion identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation.
Results
On Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolving baselines across twelve benchmarks, including methods that rely on external GPT-4o supervision, without a single regression:
- +3.60 on OCRBench
- +3.38 on MMStar Fine-Grained Perception
- +3.08 on MMStar Logical Reasoning
---
*Auto-collected from zhichai.net.*