Paper Overview
Field: Computer Vision (CV) Authors: Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar Published: 2026-08-11 arXiv: 2508.03807
Introduction
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models.
CVPD: Contrastive Counterfactual Visual Process Distillation
The authors introduce CVPD, which, to the best of their knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs.
CVPD identifies visual blind spots: regions where zooming into the region changes and sharpens the model's answer distribution, while removing the same region leaves the full-image behavior largely unchanged. Such regions reveal perceptual information that the model can encode but fails to consistently exploit under full-image conditions.
The framework:
- Proposes a three-stage counterfactual criterion to identify these regions directly from the model's own responses
- Converts them into dense contrastive supervision for self-distillation, without any external annotations, tools, or teacher models
- OCRBench: +3.60
- MMStar fine-grained perception: +3.38
- MMStar logical reasoning: +3.08
- Broader multimodal benchmarks: maintained or improved performance
- arXiv paper: <https://arxiv.org/abs/2508.03807>
Results
Evaluated on Qwen3-VL-8B-Instruct, CVPD outperforms six self-evolution baselines (including approaches that rely on external GPT-4o supervision) across 12 benchmarks, with no performance regressions: