Paper Overview
Field: Computer Vision (CV) Authors: Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar Published: 2026-08-11 arXiv: 2508.03807
Introduction
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models.
CVPD: Contrastive Counterfactual Visual Process Distillation
The paper introduces CVPD, which, to the best of the authors' knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs.
Visual Blind Spots
CVPD identifies visual blind spots—regions where:
- Zooming into the region changes and sharpens the model's answer distribution
- Removing the same region leaves the full-image behavior largely unchanged
- OCRBench: +3.60
- MMStar fine-grained perception: +3.38
- MMStar logical reasoning: +3.08
- Broader multimodal benchmarks: performance maintained or improved
- arXiv: https://arxiv.org/abs/2508.03807
These regions reveal perception information the model can encode but fails to consistently exploit under full-image conditions.
Three-Stage Counterfactual Criterion
The method uses a three-stage counterfactual criterion that identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation—no external teachers or annotations required.
Results
On Qwen3-VL-8B-Instruct, CVPD outperforms 6 self-evolution baselines across 12 benchmarks—including methods relying on external GPT-4o supervision—with no performance regressions: