Paper Overview
- Field: NLP
- Authors: Zilin Xiao, Qi Ma, Chun-cheng Jason Chen
- Published: 2025-06-13
- arXiv: 2506.10670
- Reasoning-aware retrieval: The retriever is trained to surface demonstrations by their usefulness for solving the target problem, not surface similarity.
- Diversity of retrieved contexts: The analysis shows reasoning-aware retrieval uncovers complementary solution strategies, providing different reasoning scaffolds for a single question.
- Consistent gains over RFT: On challenging mathematical reasoning benchmarks, RA-RFT consistently outperforms standard reinforcement fine-tuning methods.
- AIME 2025 results: With average@32 accuracy, RA-RFT improves over GRPO by 7.1 and 2.8 percentage points for Qwen3-1.7B and Qwen3-4B, respectively.
- Complementary dimension: The results suggest reasoning-aware retrieval is an improvement axis orthogonal to advances in reward design or training curricula.
Abstract
Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowledge, yet conventional retrieval based on lexical or semantic similarity is poorly suited for complex reasoning tasks: a semantically similar problem may demand an entirely different solution strategy, while a superficially different problem may share the same underlying reasoning pattern.
The authors propose Retrieval-Augmented Reinforcement Fine-Tuning (RA-RFT), a post-training framework that teaches language models to reason by analogy. RA-RFT uses gold-relevance distillation to train a retriever that ranks contexts by expected reasoning benefit rather than semantic overlap, and then fine-tunes the policy model via reinforcement fine-tuning with retrieved analogous demonstrations, enabling the model to exploit reasoning trajectories under verifiable outcome rewards.
Key Findings
Why It Matters
The work reframes retrieval for reasoning: instead of retrieving similar texts, models learn to find and use analogous reasoning patterns, combining the strengths of retrieval augmentation and reinforcement fine-tuning for math and other verifiable reasoning tasks.
--- *Auto-collected on 2026-06-13*