GOLF: From Scalar Rewards to Natural Language Feedback in RLHF
> Paper: Bootstrapping Exploration with Group-Level Natural Language Feedback in Reinforcement Learning > arXiv: 2603.04597 | March 6, 2026 > Institutions: Harbin Institute of Technology × Xiaohongshu > Code: https://github.com/LuckyyySTA/GOLF
One-line summary
GOLF upgrades RLHF from a learner that "only reads scores" to one that "reads critiques and summarizes experience." By aggregating external critiques and in-group attempts, it injects high-quality improvement samples adaptively in low-reward regions, achieving a 2.2x gain in sample efficiency and beating the strongest baseline by 22.7% on unverifiable tasks (per the post).
The problem: scalar rewards are too sparse
Under standard RLHF, a model may generate a refusal-like response and receive only Reward = -1, with no signal about what went wrong or how to fix it. This forces costly blind trial-and-error, and all-zero reward groups collapse group-normalized advantages, making gradients vanish.
In reality, LLMs often receive far richer supervision than a single +1/-1:
- Natural-language user feedback ("this answer is too verbose")
- Runtime errors from code execution (
IndexError: list index out of range) - Textual critiques from judge models ("step 3's assumption fails because...")
- Critique A flags a wrong assumption at step 2, critique B flags a computation error at step 3 — the model learns to check both
- If response 1's first half is sound and response 2's second half is solid, refinement can combine them
- Common failure patterns reveal systematic blind spots, so one fix benefits the whole policy
- Paper: arXiv:2603.04597
- Code: https://github.com/LuckyyySTA/GOLF
- Baselines: GRPO (Shao et al., 2024), Critique-GRPO (Zhang et al., 2025), SDPO (Hübotter et al., 2026)
These contain explicit error diagnoses, cross-attempt comparisons, and concrete revision suggestions — but traditional RLHF compresses them all into one scalar.
Core method: three mechanisms forming a virtuous cycle
1. Aggregated feedback: turning critiques into teaching material
A single failed response's critique has a narrow view; aggregating critiques across a group of failures reveals common patterns, complementary ideas, and systematic improvement directions.
Pipeline: 1. Sample N responses per prompt: G_gen(x) = {y^(1), ..., y^(N)} 2. Collect reward and critique for each: (r^(i), c^(i)) 3. Gather all failures: F(x) = {(y^(i), c^(i)) | r^(i) = 0} 4. Build an aggregated refinement prompt listing each candidate with its score and feedback, asking the model to synthesize an improved response by learning from identified mistakes, keeping strengths, and combining the best parts of all candidates 5. Sample a refinement group from the aggregated prompt: G_ref(x) = {ỹ^(j)} 6. Score refinements and keep the successful ones
Why aggregation beats single critiques:
2. Adaptive injection: help only when needed
If refinement samples are injected every time, the model becomes dependent on this crutch and loses autonomous exploration. GOLF computes the group mean reward s(x) = (1/N) Σ r(x,y) and injects only when s(x) < τ (default τ = 1/N, i.e., fewer than one success on average), replacing one failed sample with a randomly chosen successful refinement. High-reward groups are left alone so the model explores by itself.
3. Joint policy optimization
A counterintuitive finding: after standard RL fine-tuning, asking the model to self-refine at test time can *degrade* performance — RL on direct generation alone does not preserve the "read critique → revise answer" ability.
GOLF therefore collects two rollout groups per prompt — a generation group from π(·|x) and a refinement group from π(·|p_agg(x)) — merges them into a joint batch, computes advantages within each group, and updates the policy with GRPO. A key hyperparameter reshapes off-policy ratios as f(u) = u/(u+λ) with λ = 0.1, preventing off-policy samples from dominating while omitting clipping to emphasize low-probability but valid actions.
This creates a virtuous cycle: better refinement → better exploration scaffolding → more rewarded trajectories discovered → better generation → higher-quality refinements.
Results
Unverifiable tasks (conversation, writing, general ability):
| Model | Method | AlpacaEval | WildBench | ArenaHard-v2 | Avg | |---|---|---|---|---|---| | Llama-3.1-8B | Critique-GRPO | 43.31 | 25.09 | 13.73 | 40.92 | | | GOLF | 69.67 | 34.42 | 25.03 | 50.19 (+9.27) | | Qwen-3-8B | Rubric-as-Reward | 68.88 | 67.09 | 50.08 | 67.08 | | | GOLF | 71.94 | 68.16 | 52.00 | 69.26 (+2.18) |
Sample efficiency: GOLF reaches AlpacaEval performance in 80 steps vs 180 for Critique-GRPO → 2.25x.
Verifiable tasks (Qwen-3-8B): GOLF achieves AIME24 58.49, AIME25 41.65, AMC23 80.74, IFBench 38.33, IFEval 87.80, beating both GRPO and Critique-GRPO across the board.
Code generation (LCBv6 Avg@4): GRPO 44.08, SDPO 47.52, GOLF 47.71 with 1.5x sample efficiency.
Pass@k analysis shows GOLF beats GRPO across the whole k range — higher Pass@1 (better single-sample quality) and higher Pass@128 (richer diversity of successful solutions).
Why it works
1. Failures teach more than successes. Aggregating critiques of multiple failures yields more instructive improvement samples than imitating successes alone. 2. Feedback density beats precision. Scalar rewards are extremely sparse (one number per response); language feedback is dense (a paragraph per response). GOLF shows dense language feedback in RL training substantially boosts sample efficiency, challenging the assumption that RL must rely on scalar rewards. 3. Scaffolding philosophy. Like scaffolding in education, adaptive injection helps only when the learner is stuck and withdraws when it can proceed independently. 4. Joint training prevents imbalance. Training both generation and refinement keeps the policy capable on both tasks, like training offense and defense together.
Limitations and open questions
1. Dependence on judge quality: inaccurate or vague critiques weaken aggregated refinement. 2. Context length of aggregation: growing group size N makes the aggregated prompt longer; how to trade off completeness vs context budget? 3. Evaluation of unverifiable tasks: GPT-4o as judge introduces its own biases. 4. Cross-task transfer: effective on math, code, and instruction following — how does it fare on creative writing and multi-turn dialogue?