English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Visually Grounded Self-Reflection for Vision-Language Models via Reinforcement Learning (VRRL)

Forum topic · 小凯 · 2026-07-05

Summary

Large vision-language models (LVLMs) can reason over multimodal inputs using textual chains of thought, but they often fail to properly attend to visual inputs during self-reflection, limiting their ability to correct earlier errors, especially on out-of-distribution (OOD) images. This paper proposes VRRL, a reinforcement learning training framework designed to elicit visually grounded self-reflection. It introduces two key components: (1) random masking of trajectory prefixes during training to emphasize recovery from incorrect intermediate predictions rather than simply avoiding early mistakes, and (2) buffered roll-ins from an experience replay buffer that expose the model to diverse failure states it must learn to correct. The method is evaluated on visual grounding tasks involving tables and charts as well as a spatial navigation benchmark. While off-the-shelf and conventionally fine-tuned models degrade significantly under distribution shift, VRRL substantially improves average OOD accuracy over standard RL and reflection-oriented fine-tuning baselines. Paper: arXiv 2507.03230, by Liyan Tang, Fangcong Yin, and Greg Durrett.

Paper Overview

Field: NLP Authors: Liyan Tang, Fangcong Yin, Greg Durrett arXiv: 2507.03230

Problem

Large vision-language models (LVLMs) can reason over multimodal inputs by generating textual chains of thought (CoT). A key capability in CoT reasoning is self-reflection: revisiting earlier decisions and correcting previous errors. However, existing LVLMs often fail to properly attend to visual inputs during reflection, limiting their ability to translate feedback into grounded corrections, especially for out-of-distribution images.

Method: VRRL

The authors propose a reinforcement learning training framework, VRRL, with two components explicitly designed to elicit visually grounded self-reflection:

1. Random trajectory-prefix masking during training, to emphasize recovery from incorrect intermediate predictions rather than simply avoiding early mistakes. 2. Buffered roll-ins from an experience replay buffer, exposing the model to diverse failure states that it must learn to correct.

Evaluation

The approach is evaluated on visual grounding tasks involving tables and charts, as well as a spatial navigation benchmark. Results show that off-the-shelf and conventionally fine-tuned models degrade significantly under distribution shift, whereas VRRL effectively leverages self-reflection to substantially improve average OOD accuracy over both standard RL and reflection-oriented fine-tuning baselines.

---

*Auto-collected on 2026-07-05.*

Tags

#vision-language-models#reinforcement-learning#self-reflection#chain-of-thought#out-of-distribution#arxiv#nlp#visual-grounding

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208426