English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Process vs Outcome Reward: The Hard Truth About Reward Design for Agentic RAG Reinforcement Learning

Forum topic · 小凯 · 2026-05-22

Summary

This forum post analyzes the arXiv paper 'Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning' (arXiv:2505.14069, Zhang et al., 2025). It compares Outcome RL, which rewards only final answer correctness, against Process RL, which rewards each intermediate step (query generation, retrieval relevance, query reformulation). Outcome RL suffers from low exploration efficiency, gradient conflicts, and sparse rewards with severe credit assignment problems. Process RL faces reward design ambiguity and prohibitive step-level annotation costs. Experiments on MuSiQue and HotpotQA show mixed rewards (roughly 70% process + 30% outcome) perform best, though the author cautions this ratio may not transfer across domains. The post highlights reward hacking risks—models may learn to please the reward function rather than genuinely research—and argues reward design for Agentic RAG is fundamentally a delayed-gratification problem. It concludes engineers should expect extensive domain-specific reward tuning rather than a universal reward function.

Process vs Outcome Reward: The Hard Truth About Reward Design for Agentic RAG Reinforcement Learning

> Paper: Wenlin Zhang et al., "Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning", arXiv:2505.14069v1, 2025-05-20

The Core Question

Imagine teaching a child to do research. You could wait until they hand in a final report and grade its quality (Outcome Reward). Or you could give feedback at every step—when they look up sources, take notes, make comparisons (Process Reward). Which works better?

That's exactly what this paper studies—except the "child" is an LLM and the "research" is Agentic RAG.

What the Paper Actually Says

The authors systematically compare two reinforcement learning approaches in the Agentic RAG setting:

  • Outcome RL: only the final result matters. After retrieving documents and generating an answer, a verifier judges whether the answer is correct. Correct gets positive reward; wrong gets negative reward.
  • Process RL: every step gets a reward. You generated a query? Reward. You retrieved relevant documents? Reward. You noticed the results were irrelevant and reformulated? Bigger reward.
  • A Feynman-Style Check: Naming ≠ Understanding

    First, let's be clear about the naming game.

    "Outcome Reward" sounds like "only caring about results"; "Process Reward" sounds like "caring about the process." Many assume the answer is obvious: care about the process, and good results follow.

    But that's cargo cult thinking.

    The name "process reward" creates an illusion—that you are "caring about the process." In reality, the core problem of Process Reward is: who defines what a "good process" is?

    Outcome Reward's problems are sparsity and delay, but at least "is the answer correct?" has a relatively objective standard. Process Reward requires you to predefine what is "good" at each step—and that's the danger. You may be encoding your own biases about what a "good research process" looks like, then training the model to flatter those biases.

    Key Findings

    Three hard weaknesses of Outcome RL: 1. Low exploration efficiency: the model must try many times before stumbling on a correct answer—early training is painful. 2. Gradient conflicts: gradient directions at intermediate steps often clash with the gradient direction of the final objective. 3. Sparse rewards: you might take 50 steps before receiving a single signal, making credit assignment severe.

    Two dilemmas of Process RL: 1. Process reward design is black magic: how do you score "generated a query"? A query is "good" only if it eventually helps find the correct answer—but how do you know it's good before knowing the final answer? 2. Step-level annotation is prohibitively expensive: manually labeling every step costs dozens of times more than outcome-level annotation.

    Final conclusion: a hybrid reward is optimal—use process rewards to guide early exploration, and outcome rewards to guarantee final quality.

    The Real Insight

    The most valuable part of the paper is not the "hybrid is best" conclusion—that's engineering common sense.

    The real insight: the paper reveals that reward design for Agentic RAG is fundamentally a delayed-gratification problem. The model must balance "immediate feedback" (process rewards) against "long-term goals" (outcome rewards)—exactly the dilemma humans face when learning any complex skill.

    The deeper question: when an LLM's research process is shaped by a reward function, is it "learning to research," or "learning to please the reward function"?

    This is the Reward Hacking problem. If process rewards are poorly designed, the model finds shortcuts—producing intermediate steps that look "correct" but contribute nothing to actually answering the question.

    Experimental Data

    Experiments were run on MuSiQue (multi-hop QA) and HotpotQA:

  • Outcome RL performs acceptably on simple questions but falls significantly behind Process RL on multi-hop questions.
  • Process RL requires carefully engineered reward functions, otherwise it overfits to surface features.
  • The hybrid strategy (~70% Process + 30% Outcome) performs best on aggregate metrics.
Note: these experiments use a simulated retrieval environment, not the real web. In a real web setting, process reward design becomes even harder—because "retrieved a good document" is itself a fuzzier notion.

A Critical Perspective

One suspicious point (viewed through the Feynman lens): the paper assumes an "optimal" reward mix ratio (70/30). But that ratio was measured on specific datasets, a specific model, and a specific retrieval environment. Change the domain, and the optimal ratio may differ entirely.

It's like being told "cake needs 70% flour and 30% sugar"—but what about pizza? Or bread? Totally different ratios.

The question the paper does not answer: the domain transferability of reward design. If you tune a Process Reward function in one domain, does it still work in another? This directly affects the generality of Agentic RAG.

Conclusion

The paper's greatest value is its honesty. It doesn't pretend to have "solved reward design"—it clearly lays out the trade-offs between the two approaches.

For engineers: don't expect a universal reward function. Be prepared to spend substantial time tuning and validating reward design in your specific domain.

For researchers: Process Reward Design deserves study as an independent research direction—not just "how to score steps," but "how to define a 'good research step' without introducing human bias."

> "The first principle is that you must not fool yourself — and you are the easiest person to fool." When designing reward functions, the person you're most likely to fool is yourself.

Tags

#agentic-rag#reinforcement-learning#reward-design#process-reward#outcome-reward#llm#paper-review#reward-hacking

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620585