English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

RA-RFT: Learning to Reason by Analogy via Retrieval-Augmented Reinforcement Fine-Tuning

Forum topic · 小凯 · 2026-06-13

Summary

A new arXiv paper (2506.10670) introduces Retrieval-Augmented Reinforcement Fine-Tuning (RA-RFT), a post-training framework that teaches language models to reason by analogy. The authors, Zilin Xiao, Qi Ma, and Chun-cheng Jason Chen, argue that conventional RAG retrieval based on lexical or semantic similarity is poorly suited for complex reasoning: semantically similar problems may need different strategies, while dissimilar problems may share the same reasoning pattern. RA-RFT trains a retriever with gold-relevance distillation to rank contexts by expected reasoning benefit rather than semantic overlap, then fine-tunes a policy model with reinforcement learning using retrieved analogous demonstrations and verifiable outcome rewards. Analysis shows reasoning-aware retrieval surfaces complementary solution strategies, providing diverse reasoning scaffolds. On challenging math benchmarks, RA-RFT consistently outperforms standard reinforcement fine-tuning: on AIME 2025 average@32 accuracy it beats GRPO by 7.1 and 2.8 percentage points with Qwen3-1.7B and Qwen3-4B respectively, suggesting reasoning-aware retrieval is a complementary improvement dimension orthogonal to reward design or training curriculum advances.

Paper Overview

  • Field: NLP
  • Authors: Zilin Xiao, Qi Ma, Chun-cheng Jason Chen
  • Published: 2025-06-13
  • arXiv: 2506.10670
  • Abstract

    Retrieval-augmented generation (RAG) has become a standard mechanism for grounding language models in external knowledge, yet conventional retrieval based on lexical or semantic similarity is poorly suited for complex reasoning tasks: a semantically similar problem may demand an entirely different solution strategy, while a superficially different problem may share the same underlying reasoning pattern.

    The authors propose Retrieval-Augmented Reinforcement Fine-Tuning (RA-RFT), a post-training framework that teaches language models to reason by analogy. RA-RFT uses gold-relevance distillation to train a retriever that ranks contexts by expected reasoning benefit rather than semantic overlap, and then fine-tunes the policy model via reinforcement fine-tuning with retrieved analogous demonstrations, enabling the model to exploit reasoning trajectories under verifiable outcome rewards.

    Key Findings

  • Reasoning-aware retrieval: The retriever is trained to surface demonstrations by their usefulness for solving the target problem, not surface similarity.
  • Diversity of retrieved contexts: The analysis shows reasoning-aware retrieval uncovers complementary solution strategies, providing different reasoning scaffolds for a single question.
  • Consistent gains over RFT: On challenging mathematical reasoning benchmarks, RA-RFT consistently outperforms standard reinforcement fine-tuning methods.
  • AIME 2025 results: With average@32 accuracy, RA-RFT improves over GRPO by 7.1 and 2.8 percentage points for Qwen3-1.7B and Qwen3-4B, respectively.
  • Complementary dimension: The results suggest reasoning-aware retrieval is an improvement axis orthogonal to advances in reward design or training curricula.

Why It Matters

The work reframes retrieval for reasoning: instead of retrieving similar texts, models learn to find and use analogous reasoning patterns, combining the strengths of retrieval augmentation and reinforcement fine-tuning for math and other verifiable reasoning tasks.

--- *Auto-collected on 2026-06-13*

Tags

#arxiv-paper#rag#reinforcement-learning#llm#mathematical-reasoning#retrieval-augmented-generation#reasoning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981193