English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

KuaiRP Series Role-playing Models: Technical Report

Forum topic · 小凯 · 2026-09-13

Summary

A technical report introducing the KuaiRP series of role-playing models from Kuaishou researchers (arXiv 2609.11127). The work targets four goals for dedicated role-playing models: simplified prompt engineering, highly stable output quality, built-in domain world knowledge, and efficient deployment at small parameter size. Because injecting deep domain knowledge typically causes catastrophic forgetting of general agent capabilities, the authors propose a multi-stage training pipeline: (1) a standardized character template with an SFT data pipeline built on user behavior simulation and reverse profile filtering; (2) a reinforcement learning stage with a rule-based composite reward function that eliminates degenerations such as length expansion and repetitive generation; and (3) a novel self-distillation paradigm using Two-stage On-Policy Distillation (OPD) with Cumulative-Divergence Decay (CDD), where the domain-adapted model serves as teacher and the original base model as student. Experiments show KuaiRP models match state-of-the-art proprietary models in role-playing fidelity while recovering general agent capabilities at very low deployment cost.

Paper Overview

  • Research areas: cs.AI, cs.CL
  • Authors: Yipeng Wang, Ziwei Zhang, Jiahui Zhang, Qi Gan, Kai Sheng
  • Published: 2026-09-13
  • arXiv: 2609.11127

Abstract

This paper introduces the complete technical solution for the KuaiRP series of role-playing models. The authors aim to achieve four core objectives for a dedicated role-playing model: simplified prompt engineering, highly stable output quality, built-in domain world knowledge, and high-efficiency deployment with a small parameter size.

However, effectively injecting deep domain knowledge often leads to severe catastrophic forgetting of the model's general agent capabilities. To overcome this trade-off, the paper proposes a multi-stage training pipeline:

1. SFT stage: A standardized character template is designed, and an SFT data pipeline is constructed based on user behavior simulation and reverse profile filtering. 2. RL stage: A rule-based composite reward function is used during Reinforcement Learning to eliminate common degradation phenomena such as length expansion and repetitive generation. 3. Self-distillation stage: To recover general capabilities compromised during SFT and RL, a novel self-distillation paradigm is proposed using Two-stage On-Policy Distillation (OPD) equipped with Cumulative-Divergence Decay (CDD). The domain-adapted model serves as the teacher and the original base model as the student, balancing deep domain knowledge injection with preservation of general agent capabilities.

Results

Experimental results demonstrate that the KuaiRP models not only match current state-of-the-art proprietary models in role-playing fidelity within the target domains, but also successfully recover general agent capabilities while maintaining extremely low deployment costs.

---

*Auto-collected on 2026-09-13*

Tags

#role-playing#llm#reinforcement-learning#distillation#catastrophic-forgetting#arxiv#ai#chatbot

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634798