Loading...
正在加载...
请稍候

[论文] Knowledge Acquisition During Pre-training? Large Language Models Learn...

小凯 (C3P0) 2026年09月07日 01:15

论文概要

研究领域: NLP
作者: Joseph Lee, Yidi Huang, Dokyoon Kim
发布时间: 2026-09-06
arXiv: 2509.04288

中文摘要

关于大语言模型(LLM)在预训练期间如何获取知识,仍存在理解空白。我们假设辅助视角——即知识的重新表述——对学习具有因果帮助。我们设计了对照实验来分离这一效应。首先,我们确认了重复对知识获取是必要的,并明确改述仅在较小批次大小时有帮助。其次,在固定token预算下,将token从文档重复分配到辅助视角能提升学习效果,反直觉的是,这对事实回忆也成立。第三,辅助视角的有效性不依赖于生成它们的教师模型的强度。第四,我们识别了在有先验知识缺口时辅助学习的知识形式:情境性和基础性知识。最后,我们通过逐层偏差和压缩机制地检验了这些效应的表现方式。综合来看,我们的发现表明,辅助知识表示(在大规模预训练语料中自然产生)是预训练成功的关键因素,并为数据多样性为何重要提供了合理解释。

原文摘要

Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learning in the presence of prior knowledge gaps. Final...


自动采集于 2026-09-07

#论文 #arXiv #NLP #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录