English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RiVER: Reinforcement Learning Without Ground-Truth Solutions Can Improve LLMs

Forum topic · 小凯 · 2026-06-27

Summary

Reinforcement learning with verifiable rewards (RLVR) for training large language models typically relies on ground-truth answers to assign rewards, restricting its use to tasks where correct solutions are known in advance. This paper introduces RiVER (Ranking-induced VERifiable framework), an approach that trains LLMs on score-based optimization tasks without ground-truth solutions, using deterministic execution feedback as continuous-valued supervision. By applying group-relative reinforcement learning to these continuous signals, RiVER extends RLVR to a broader class of problems where exact answers are unavailable. The paper is authored by Yingyu Lin, Qiyue Gao, and Nikki Lijing Kuang, and is available on arXiv as 2606.27369.

Overview

Field: Machine Learning Authors: Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang Published: 2026-06-27 arXiv: 2606.27369

Abstract

Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically relies on ground-truth answers to assign rewards, limiting applicability to tasks where the ground-truth solution is unknown. The authors introduce a Ranking-induced VERifiable framework (RiVER) that trains LLMs on score-based optimization tasks without ground-truth solutions, using deterministic execution feedback as continuous-valued supervision. When applying group-relative RL to such continuous feedback, RiVER extends RLVR-style training to settings without known correct answers.

Key Points

  • Standard RLVR requires ground-truth answers for reward assignment, which excludes many real-world optimization tasks.
  • RiVER replaces ground-truth answers with deterministic execution feedback, treated as continuous-valued supervision.
  • A ranking-induced, group-relative RL scheme is applied to these continuous signals.
  • The framework broadens the scope of verifiable-reward reinforcement learning for LLM training.
*Auto-collected on 2026-06-27*

Tags

#reinforcement-learning#llm#rlvr#machine-learning#arxiv#training-methods#verifiable-rewards

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178208173