[论文] Base Models Can Reason By Taking a Cue From Training Data
论文概要 研究领域: LLM 作者: Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros 发布时间: 2026-10-05 arXiv: 2610.06851
论文概要
研究领域: LLM 作者: Sophie L. Wang, Amil Dravid, Rulin Shao, Kevin Farhat, Sewon Min, Alexei A. Efros 发布时间: 2026-10-05 arXiv: 2610.06851中文摘要
本文研究训练数据如何在基座模型回复的起始 token 与后续推理行为之间建立关联。首先,我们证明固定特定的起始 token 线索可以让基座模型的表现与其经过强化学习(RL)训练的对应版本相媲美。例如,线索 ".\n\nOkay" 将 Olmo-3-7B 在 MATH-500 上 pass@1 准确率从42%提升至78%,而 "Alright," 将 Qwen3-14B 从72%提升至87%。其次,RL 使这些线索更可能出现,而固定这些线索可以恢复 RL 相对于基座模型的大部分性能增益。第三,我们通过因果数据干预将任意一个词(如"chicken")转化为有效的推理线索,或移除已有线索的效果。类似的编辑使提示指令 "Think duck duck goose" 与 "Think step by step" 在激发推理方面同样有效。不同线索诱导的隐状态表示与训练集中不同类型的文档相关。最后,我们在语言模型安全性案例研究中发现,不同线索会引发不同的拒绝和遵从行为,对应不同类型的训练数据。原文摘要
We study how training data creates associations between starting token cues and reasoning behavior in base models. Fixing particular cues makes base models competitive with RL-trained counterparts. Causal data interventions can turn arbitrary words into reasoning cues. Different cues elicit distinct refusal and compliance behaviors corresponding to different training data types.*自动采集于 2026-10-07*
#论文 #arXiv #LLM #小凯