Process vs Outcome Reward: The Hard Truth About Reward Design for Agentic RAG Reinforcement Learning
> Paper: Wenlin Zhang et al., "Process vs. Outcome Reward: Which is Better for Agentic RAG Reinforcement Learning", arXiv:2505.14069v1, 2025-05-20
The Core Question
Imagine teaching a child to do research. You could wait until they hand in a final report and grade its quality (Outcome Reward). Or you could give feedback at every step—when they look up sources, take notes, make comparisons (Process Reward). Which works better?
That's exactly what this paper studies—except the "child" is an LLM and the "research" is Agentic RAG.
What the Paper Actually Says
The authors systematically compare two reinforcement learning approaches in the Agentic RAG setting:
- Outcome RL: only the final result matters. After retrieving documents and generating an answer, a verifier judges whether the answer is correct. Correct gets positive reward; wrong gets negative reward.
- Process RL: every step gets a reward. You generated a query? Reward. You retrieved relevant documents? Reward. You noticed the results were irrelevant and reformulated? Bigger reward.
- Outcome RL performs acceptably on simple questions but falls significantly behind Process RL on multi-hop questions.
- Process RL requires carefully engineered reward functions, otherwise it overfits to surface features.
- The hybrid strategy (~70% Process + 30% Outcome) performs best on aggregate metrics.
A Feynman-Style Check: Naming ≠ Understanding
First, let's be clear about the naming game.
"Outcome Reward" sounds like "only caring about results"; "Process Reward" sounds like "caring about the process." Many assume the answer is obvious: care about the process, and good results follow.
But that's cargo cult thinking.
The name "process reward" creates an illusion—that you are "caring about the process." In reality, the core problem of Process Reward is: who defines what a "good process" is?
Outcome Reward's problems are sparsity and delay, but at least "is the answer correct?" has a relatively objective standard. Process Reward requires you to predefine what is "good" at each step—and that's the danger. You may be encoding your own biases about what a "good research process" looks like, then training the model to flatter those biases.
Key Findings
Three hard weaknesses of Outcome RL: 1. Low exploration efficiency: the model must try many times before stumbling on a correct answer—early training is painful. 2. Gradient conflicts: gradient directions at intermediate steps often clash with the gradient direction of the final objective. 3. Sparse rewards: you might take 50 steps before receiving a single signal, making credit assignment severe.
Two dilemmas of Process RL: 1. Process reward design is black magic: how do you score "generated a query"? A query is "good" only if it eventually helps find the correct answer—but how do you know it's good before knowing the final answer? 2. Step-level annotation is prohibitively expensive: manually labeling every step costs dozens of times more than outcome-level annotation.
Final conclusion: a hybrid reward is optimal—use process rewards to guide early exploration, and outcome rewards to guarantee final quality.
The Real Insight
The most valuable part of the paper is not the "hybrid is best" conclusion—that's engineering common sense.
The real insight: the paper reveals that reward design for Agentic RAG is fundamentally a delayed-gratification problem. The model must balance "immediate feedback" (process rewards) against "long-term goals" (outcome rewards)—exactly the dilemma humans face when learning any complex skill.
The deeper question: when an LLM's research process is shaped by a reward function, is it "learning to research," or "learning to please the reward function"?
This is the Reward Hacking problem. If process rewards are poorly designed, the model finds shortcuts—producing intermediate steps that look "correct" but contribute nothing to actually answering the question.
Experimental Data
Experiments were run on MuSiQue (multi-hop QA) and HotpotQA:
A Critical Perspective
One suspicious point (viewed through the Feynman lens): the paper assumes an "optimal" reward mix ratio (70/30). But that ratio was measured on specific datasets, a specific model, and a specific retrieval environment. Change the domain, and the optimal ratio may differ entirely.
It's like being told "cake needs 70% flour and 30% sugar"—but what about pizza? Or bread? Totally different ratios.
The question the paper does not answer: the domain transferability of reward design. If you tune a Process Reward function in one domain, does it still work in another? This directly affects the generality of Agentic RAG.
Conclusion
The paper's greatest value is its honesty. It doesn't pretend to have "solved reward design"—it clearly lays out the trade-offs between the two approaches.
For engineers: don't expect a universal reward function. Be prepared to spend substantial time tuning and validating reward design in your specific domain.
For researchers: Process Reward Design deserves study as an independent research direction—not just "how to score steps," but "how to define a 'good research step' without introducing human bias."
> "The first principle is that you must not fool yourself — and you are the easiest person to fool." When designing reward functions, the person you're most likely to fool is yourself.