Paper Overview
Field: NLP Authors: Ryan Bahlous-Boldi, Isha Puri, Idan Shenfeld Published: 2025-05-23 arXiv: 2505.17385
Abstract (English Translation)
Language models now need to generalize to new environments out of the box and operate within inference-time search procedures, such as AlphaEvolve, which use a variety of task-specific reward functions to select rollouts. Unfortunately, standard LLM post-training paradigms optimize pre-specified scalar rewards, which tends to produce low-entropy response distributions—making it difficult for current LLMs to exhibit the diversity required for inference-time search.
The authors propose Vector Policy Optimization (VPO), a reinforcement learning algorithm that explicitly trains policies to anticipate diverse downstream reward functions and produce diverse solutions. VPO leverages the fact that rewards in practice are often vector-valued—for example, correctness on each test case in code generation, or multiple distinct user personas and reward models.
VPO can serve as a drop-in replacement for the GRPO advantage estimator, but instead trains the LLM to output a *set* of solutions, where each solution specializes in a different trade-off in the vector reward space.
Key Results
- Across four tasks, VPO matches or beats the strongest scalar RL baselines on test-time search (e.g., pass@k and best@k).
- The gap grows larger with bigger search budgets.
- For evolutionary search, VPO-trained models unlock problems that GRPO models completely fail to solve.
- As test-time search becomes more standardized, optimizing for diversity may become the default post-training objective.
*Auto-collected on 2026-05-23*