English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

CORA: Bridging the Thinking-Answer Gap in Multimodal RLVR

Forum topic · 小凯 · 2026-06-16

Summary

CORA (Consistency-Oriented Reasoning Alignment) is a method for reducing thinking-answer inconsistency in multimodal reinforcement learning with verifiable rewards (RLVR). The authors—Jiayue Cao, Zhicong Lu, and Xuehan Sun—analyze rollouts collected during GRPO training and post-RLVR evaluation outputs of large vision-language models (LVLMs), showing that semantic inconsistency between the reasoning process and the final answer persists throughout training and remains present at inference. While prior RLVR work focuses on improving visual coverage of reasoning traces and mitigating visual hallucinations, it underestimates this reasoning-answer mismatch. CORA introduces thinking-answer semantic consistency into RLVR via a lightweight, plug-and-play consistency reward model, combined with Hybrid Reward Advantage Segmentation (HRAS) to stably coordinate task and consistency optimization. Experiments on representative multimodal reasoning benchmarks and mainstream LVLMs show that CORA improves task performance while alleviating thinking-answer inconsistency, producing more faithful reasoning traces. Paper: arXiv:2606.14691.

Paper Overview

  • Field: Computation and Language (CL)
  • Authors: Jiayue Cao, Zhicong Lu, Xuehan Sun
  • arXiv: 2606.14691
  • Problem

    Reinforcement learning with verifiable rewards (RLVR) has successfully elicited reasoning capabilities in large language models, motivating its extension to multimodal scenarios. Existing methods mainly focus on improving the visual coverage of reasoning traces and mitigating visual hallucinations, but underestimate the semantic inconsistency between the reasoning process and the final answer.

    Analysis

    The authors study thinking-answer inconsistency in RLVR for large vision-language models (LVLMs) through thorough analyses of:

  • Rollouts collected throughout the Group Relative Policy Optimization (GRPO) training process
  • Post-RLVR evaluation outputs
  • The findings show this issue persists during training and remains present during inference.

    Method: CORA

    Motivated by the analysis, CORA (Consistency-Oriented Reasoning Alignment) consists of:

  • A lightweight, plug-and-play consistency reward model that brings thinking-answer semantic consistency into RLVR
  • Hybrid Reward Advantage Segmentation (HRAS) to stably coordinate task and consistency optimization
  • Results

    Extensive experiments on representative multimodal reasoning benchmarks and mainstream LVLMs show that CORA:

  • Improves task performance
  • Effectively mitigates thinking-answer inconsistency
  • Produces more faithful reasoning traces

Original Abstract (partial)

> Reinforcement learning with verifiable rewards (RLVR) has successfully elicited the reasoning capabilities of large language models, motivating its extension to multimodal scenarios. Existing methods primarily focus on improving the visual coverage of reasoning traces and mitigating visual hallucinations, but underestimate the semantic inconsistency between the reasoning process and the final answer...

Source: arXiv:2606.14691. Auto-collected on 2026-06-16.

Tags

#multimodal-reasoning#rlvr#large-vision-language-models#grpo#reinforcement-learning#consistency#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981385