[论文] LittleLearner: Language Models Under Pedagogically Controlled Knowledg...
论文概要
研究领域: NLP 作者: Fanfei Li, Jana Zeller, Manuel Prada-Corral, Thaddäus Wiedemer, Prasanna Mayilvahanan, Ryan Cotterell, Wieland Brendel 发布时间: 2026-08-13 arXiv: 2608.13545中文摘要
现代语言模型在异构的网页规模文本语料库上训练。因此,研究知识和技能获取很困难,因为事先接触相关内容的情况难以表征。为解决这一挑战,我们引入LITTLECURRICULUM,一个精心策划的880亿token预训练语料库,专门针对美国小学教材,明确排除五年级以上教授的概念、事实和词汇。在LITTLECURRICULUM上从头训练一个50亿参数的LLM得到LITTLELEARNER,该模型具有足够的语言能力进行开放式评估,但具有与可解释课程指南映射的明确知识和能力边界。我们发布LITTLECURRICULUM和LITTLELEARNER作为一个发展受限的沙盒,用于研究模型如何在明确定义的训练范围内获取、表示和使用数据。我们在通过后训练和上下文学习注入新知识的第一组实验中展示了沙盒的效用。这些方法让LITTLELEARNER更好地利用现有知识,但不会提升超出范围的能力。我们的发现强调了这种受控环境对未来研究的价值。原文摘要
Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LITTLECURRICULUM, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LITTLECURRICULUM yields LITTLELEARNER, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LITTLECURRICULUM and LITTLELEARNER as a developmentally restricted sandbox to study how models acquire, represent, and use data under a well-defined training scope. We illustrate the sandbox's utility in a first suite of experiments on injecting new knowledge through post-training and in-context learning. These methods let LITTLELEARNER better utilize existing knowledge,但不会提升超出范围的能力. Our findings underscore the value of this controlled environment for future investigations.--- *自动采集于 2026-08-15*
#论文 #arXiv #NLP #小凯