English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Vector Policy Optimization: Training for Diversity Improves Test-Time Search in LLMs

Forum topic · 小凯 · 2026-05-23

Summary

Vector Policy Optimization (VPO) is a reinforcement learning algorithm proposed by Ryan Bahlous-Boldi, Isha Puri, and Idan Shenfeld (arXiv:2505.17385, May 2025) that trains language models to produce diverse solutions optimized for downstream reward functions. Standard LLM post-training optimizes a single scalar reward, yielding low-entropy output distributions that struggle to support test-time search methods such as best-of-k sampling or evolutionary search. VPO exploits the fact that rewards in practice are often vector-valued (e.g., correctness per test case in code generation, or multiple user personas and reward models) and serves as a drop-in replacement for the GRPO advantage estimator. Instead of one solution, VPO trains the model to output a set of solutions, each specializing in a different trade-off within the vector reward space. Across four tasks, VPO matches or beats the strongest scalar RL baselines on test-time search metrics like pass@k and best@k, with the gap widening at larger search budgets. For evolutionary search, VPO-trained models solve problems that GRPO models cannot solve at all. The authors argue that optimizing for diversity may become a default post-training objective as test-time search becomes standardized.

Paper Overview

Field: NLP Authors: Ryan Bahlous-Boldi, Isha Puri, Idan Shenfeld Published: 2025-05-23 arXiv: 2505.17385

Abstract (English Translation)

Language models now need to generalize to new environments out of the box and operate within inference-time search procedures, such as AlphaEvolve, which use a variety of task-specific reward functions to select rollouts. Unfortunately, standard LLM post-training paradigms optimize pre-specified scalar rewards, which tends to produce low-entropy response distributions—making it difficult for current LLMs to exhibit the diversity required for inference-time search.

The authors propose Vector Policy Optimization (VPO), a reinforcement learning algorithm that explicitly trains policies to anticipate diverse downstream reward functions and produce diverse solutions. VPO leverages the fact that rewards in practice are often vector-valued—for example, correctness on each test case in code generation, or multiple distinct user personas and reward models.

VPO can serve as a drop-in replacement for the GRPO advantage estimator, but instead trains the LLM to output a *set* of solutions, where each solution specializes in a different trade-off in the vector reward space.

Key Results

  • Across four tasks, VPO matches or beats the strongest scalar RL baselines on test-time search (e.g., pass@k and best@k).
  • The gap grows larger with bigger search budgets.
  • For evolutionary search, VPO-trained models unlock problems that GRPO models completely fail to solve.
  • As test-time search becomes more standardized, optimizing for diversity may become the default post-training objective.
---

*Auto-collected on 2026-05-23*

Tags

#reinforcement-learning#llm#vector-policy-optimization#test-time-search#grpo#rlhf#diversity#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620658