Paper Overview
- Field: Computer Vision (CV)
- Authors: Shravan Venkatraman, Omkar Thawakar, Ritesh Thawkar
- Published: 2026-08-11
- arXiv: 2508.03807
- Zooming into the region changes and sharpens the model's answer distribution, while
- Removing the same region leaves the full-image behavior largely unchanged.
- Outperforms 6 self-evolution baselines (including methods relying on external GPT-4o supervision) across 12 benchmarks, with no performance regressions.
- Gains +3.60 on OCRBench.
- Gains +3.38 on MMStar fine-grained perception.
- Gains +3.08 on MMStar logical reasoning.
- Maintains or improves performance on broader multimodal benchmarks.
- arXiv: <https://arxiv.org/abs/2508.03807>
Abstract
Self-improvement for multimodal large language models (MLLMs) is typically driven by reward-based methods that provide only coarse scalar feedback. Distillation offers a richer alternative through dense token-level supervision, but in the visual domain it usually depends on privileged context constructed using external annotations and tools, or stronger models.
The authors introduce CVPD (Contrastive Counterfactual Visual Process Distillation), which, to the best of their knowledge, is the first fully self-contained framework for dense, on-policy, token-level visual self-distillation for MLLMs.
Key Idea: Visual Blind Spots
CVPD identifies *visual blind spots*—regions where:
Such regions reveal perceptual information that the model can encode but fails to consistently exploit under full-image conditions.
Method
A three-stage counterfactual criterion identifies these regions directly from the model's own responses and converts them into dense contrastive supervision for self-distillation—no external teachers, annotations, or tools required.
Results
Evaluated on Qwen3-VL-8B-Instruct, CVPD: