论文概要
研究领域: LLM
作者: Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros
发布时间: 2026-10-05
arXiv: 2610.06851
中文摘要
本文研究训练数据如何在基座模型回复的起始 token 与后续推理行为之间建立关联。首先,我们证明固定特定的起始 token 线索可以让基座模型的表现与其经过强化学习(RL)训练的对应版本相媲美。例如,线索 ".\n\nOkay" 将 Olmo-3-7B 在 MATH-500 上 pass@1 准确率从42%提升至78%,而 "Alright," 将 Qwen3-14B 从72%提升至87%。其次,RL 使这些线索更可能出现,而固定这些线索可以恢复 RL 相对于基座模型的大部分性能增益。第三,我们通过因果数据干预将任意一个词(如"chicken")转化为有效的推理线索,或移除已有线索的效果。类似的编辑使提示指令 "Think duck duck goose" 与 "Think step by step" 在激发推理方面同样有效。不同线索诱导的隐状态表示与训练集中不同类型的文档相关。最后,我们在语言模型安全性案例研究中发现,不同线索会引发不同的拒绝和遵从行为,对应不同类型的训练数据。
原文摘要
We study how training data creates associations between starting token cues and reasoning behavior in base models. Fixing particular cues makes base models competitive with RL-trained counterparts. Causal data interventions can turn arbitrary words into reasoning cues. Different cues elicit distinct refusal and compliance behaviors corresponding to different training data types.
自动采集于 2026-10-07
#论文 #arXiv #LLM #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。