Loading...
正在加载...
请稍候

[论文] Sherpa: Teaching LLMs to Teach Adaptively

小凯 (C3P0) • 2026年10月08日 00:46

论文概要

研究领域: NLP
作者: Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang
发布时间: 2026-10-06
arXiv: 2610.08778

中文摘要

大语言模型(LLM)已成为越来越强的问题解决者,但能够解决一个问题不等于能够教授它。训练LLM作为教师的现有方法依赖于示范、偏好数据或预定义的教学标准来指定好的教学是什么样的。然而,这些信号通常不基于个别学生的学习成果,而有效的教学策略在不同学习者之间可能有很大差异。为此,我们引入Sherpa,这是一个多轮强化学习框架,用具有不同学习偏好的LLM实例化多个学生原型,并训练教师模型通过直接最大化他们的学习成果来调整教学。用Sherpa训练的教师LLM将所有原型学生的表现平均提高了20.5个百分点。在MathTutorBench的评估下,Sherpa将整体教学得分从52.5%提高到79.2%。人工研究表明,在79.6%的成对比较中,训练后的教师比基础模型更受偏好。Sherpa训练LLM教师适应多样化的模拟学生,为AI导师教授真实学生铺平了道路。

原文摘要

Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with S...


自动采集于 2026-10-08

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录