Paper Overview
Field: Machine Learning Authors: Yingyu Lin, Qiyue Gao, Nikki Lijing Kuang Published: 2026-06-27 arXiv: 2606.27369
Chinese Abstract (translated)
Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically relies on ground-truth answers to assign rewards, limiting its applicability to tasks where the ground-truth solution is unknown. The authors introduce a Ranking-induced VERifiable framework (RiVER) that trains LLMs on score-based optimization tasks without ground-truth solutions, using deterministic execution feedback as continuous-valued supervision.
Original Abstract
Reinforcement learning with verifiable rewards (RLVR) for training LLMs typically rely on ground-truth answers to assign rewards, limiting their applicability to tasks where the ground-truth solution is unknown. We introduce a Ranking-induced VERifiable framework (RiVER) that trains LLMs on score-based optimization tasks without ground-truth solutions, using deterministic execution feedback as continuous-valued supervision. When applying group-relative RL to such conti...
*(The abstract is truncated in the source post; see the arXiv link above for the full version.)*
---
*Automatically collected on 2026-06-27.*