English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Why RL-Trained Models Merge Better Than SFT Models: Task Conflict Explained

Forum topic · ✨步子哥 · 2026-08-03

Summary

A recent arXiv paper (2607.22039) reveals a robust, counterintuitive finding in LLM model merging: models fine-tuned with reinforcement learning (RL) suffer far less performance loss when merged than models trained with supervised fine-tuning (SFT). Using Llama-3.2-3B, Llama-3.1-8B, and Mistral-Small-3-24B as base models, and training on five tasks (math, code, instruction following, logic puzzles, ranking) with PPO, GRPO, and Reinforce++, the authors show the result holds across merging methods (linear averaging, task arithmetic, TIES), RL algorithms, and base models. Three mechanisms explain it: RL's on-policy data limits gradient magnitude, its expected-return objective is milder than SFT's negative log-likelihood winner-take-all target, and RL optimizes positive and negative samples symmetrically. A theoretical conflict norm analysis shows RL's advantage function baseline cancels opposing gradients, reducing task conflict. Practical takeaway: prefer RL before model merging; SFT builds capability, RL reduces conflicts.

> 📌 This is a GEO-optimized version of the original topic — restructured with question-driven framing, structured data, and FAQ for better AI citation.

> One-line conclusion: RL-trained models merge far more gracefully than SFT-trained models, and this holds across merging methods, RL algorithms, and base models.

Why RL-Trained Models Are More "Merge-Friendly": The Fundamental Difference Between SFT and RL in Model Merging

> Paper: *Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts* > arXiv: 2607.22039 (July 24, 2025)

An Engineering Problem

Suppose you have three models: one good at math, one at code, one at instruction following — and you want to merge them into a single "all-rounder."

Model merging is a popular LLM technique: instead of retraining, you average or combine the parameters of multiple models to get a multi-task model. Sounds great.

In practice, merged models often perform *worse* than any individual source model. The culprit is task conflict: the parameter update direction of the math model may oppose that of the code model, and merging drags both capabilities down.

This paper finds a counterintuitive phenomenon: RL-trained models lose far less performance after merging than SFT-trained models. This holds regardless of merging method, RL algorithm, or base model.

Experimental Setup

The authors used three base models (Llama-3.2-3B, Llama-3.1-8B, Mistral-Small-3-24B), each trained with SFT and three RL algorithms (PPO, GRPO, Reinforce++) on five tasks:

  • Math: OpenMathInstruct-2 → GSM8K, MATH-500
  • Code: OpenCodeInstruct → HumanEval, MBPP
  • Instruction following: Tulu-3-SFT → IFEval, LiveBench
  • Logic puzzles: Knights and Knaves
  • Ranking: a ranking task
  • Models were merged pairwise, three-way, and four-way, and performance retention was measured.

    Core Finding: RL Models Lose Less After Merging

    Across all merging methods (linear averaging, task arithmetic, TIES, etc.), RL-trained models retain significantly more performance after merging than SFT-trained models. The result is independent of merging method, RL algorithm, and base model — a robust regularity, not a lucky configuration.

    Three Reasons RL Is More "Merge-Friendly"

    1. On-policy data limits gradient magnitude

    SFT uses fixed labeled data, pushing parameters hard in one direction — large updates that can overwrite other tasks' knowledge. RL uses on-policy data: the model generates its own answers, evaluated by reward signals. Updates are small corrections only where the current policy underperforms, so they are less likely to overwrite other capabilities.

    Analogy: SFT is rote memorization of standard answers, creating strong path dependence; RL is correcting mistakes from feedback, fixing only what's wrong.

    2. RL's optimization objective is inherently milder

    SFT's negative log-likelihood loss is winner-take-all: maximize the probability of the correct answer, minimize everything else — parameters "charge" in one direction. RL maximizes expected return: it only requires good answers to be *more likely* than bad ones. This mildness makes RL parameter updates more dispersed, reducing inter-task conflict.

    3. Symmetric optimization of positive and negative samples

    RL increases the probability of high-reward answers *and* decreases that of low-reward answers — a balanced, symmetric adjustment. SFT only has positive samples (gold answers) with no mechanism for pushing down bad answers, making its updates more lopsided.

    Theoretical Analysis: The Conflict Norm

    The authors quantify task conflict with a conflict norm measuring how much two tasks' parameter update directions oppose each other. Theory shows RL yields a smaller conflict norm: the advantage function's baseline (subtracting the value estimate) makes gradients from positive and negative samples cancel in expectation, leaving only "truly useful" update directions. SFT has no such cancellation, so its updates scatter and conflict more.

    Engineering Implications

    1. If you plan to merge models, prefer RL training. SFT-trained models degrade when merged; RL-trained models largely retain performance. 2. RL's value isn't "stronger" — it's more compatible. In multi-task settings, RL's worth lies not in single-task performance (often comparable to SFT) but in merge-friendliness. 3. Combine SFT and RL: SFT builds foundational capability; RL serves as de-confliction before merging.

    Honest Assessment

  • Limited scale: three base models, five tasks. Real-world multi-task settings may involve dozens of tasks with more complex conflict patterns.
  • Strong theoretical assumptions: the conflict norm derivation assumes independent parameter updates, while different layers' updates are highly correlated in practice.
  • The title — "Enough is as good as a feast" — hints at a deeper point: RL's mildness is not a defect but an advantage. In merging scenarios you don't need to learn the deepest, you need to learn the most compatibly. Inspiring, though the paper doesn't fully establish universality.
  • One-Line Summary

    RL-trained models merge better because their gradient updates are smaller, their objective milder, and their positive/negative optimization symmetric — "contentment" here isn't a moral lesson, it's a mathematical property.

    ---

    Paper: https://arxiv.org/abs/2607.22039 HTML version: https://arxiv.org/html/2607.22039v1 Open-source code: not yet released (training used verl and OpenRLHF)

    FAQ

    Q1: Who is this for?

    Practitioners, researchers, and students in AI, machine learning, and deep learning.

    Q2: What are the key takeaways?

  • An engineering problem (task conflict in model merging)
  • A controlled experimental setup across three base models, five tasks, and three RL algorithms
  • Core finding: RL-trained models retain far more performance after merging than SFT-trained models
Q3: Is there open-source code?

No official release yet; see the links above for the training frameworks used.

Tags

#model-merging#reinforcement-learning#sft#llm#task-conflict#ppo#grpo#multi-task-learning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178503907