Paper Overview
Field: NLP Authors: Liyan Tang, Fangcong Yin, Greg Durrett arXiv: 2507.03230
Problem
Large vision-language models (LVLMs) can reason over multimodal inputs by generating textual chains of thought (CoT). A key capability in CoT reasoning is self-reflection: revisiting earlier decisions and correcting previous errors. However, existing LVLMs often fail to properly attend to visual inputs during reflection, limiting their ability to translate feedback into grounded corrections, especially for out-of-distribution images.
Method: VRRL
The authors propose a reinforcement learning training framework, VRRL, with two components explicitly designed to elicit visually grounded self-reflection:
1. Random trajectory-prefix masking during training, to emphasize recovery from incorrect intermediate predictions rather than simply avoiding early mistakes. 2. Buffered roll-ins from an experience replay buffer, exposing the model to diverse failure states that it must learn to correct.
Evaluation
The approach is evaluated on visual grounding tasks involving tables and charts, as well as a spatial navigation benchmark. Results show that off-the-shelf and conventionally fine-tuned models degrade significantly under distribution shift, whereas VRRL effectively leverages self-reflection to substantially improve average OOD accuracy over both standard RL and reflection-oriented fine-tuning baselines.
---
*Auto-collected on 2026-07-05.*