[论文] Telescopic Language Models

研究领域: NLP 作者: Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Siyu Hong, Yixiao Liu, Adam…

论文概要

研究领域: NLP 作者: Zhilin Guo, Boqiao Zhang, Hakan Aktas, Kyle Fogarty, Nursena Koprucu Aslan, Wenzhao Li, Canberk Baykal, Albert Miao, Siyu Hong, Yixiao Liu, Adam Wu, Ashish Kumar Singh, Sakar Khattar, Chenliang Zhou, Weihao Xia, Cristina Nader Vasconcelos, Cengiz Oztireli 发布时间: 2026-09-28 arXiv: 2609.35769

中文摘要

一个已部署的语言模型通常需要服务多种算力预算,但为每个预算点服务仍意味着每次都要单独训练或压缩。我们训练了望远镜式语言模型(TLM)来作为连续体:一种通过随机前缀监督与全容量锚点进行监督的嵌套容量 Transformer。每一步中,容量轴上一个随机截断的前缀在全量下一 token 目标下训练,同时进行一次全容量前向-反向传播,因此训练产物在每一层深度上都是有效的语言模型。每步仅需两次前向-反向传播,无需架构改动,推理时零额外开销。固定出口套件(如 Matryoshka LM 套件)只是该设计空间中的一个点,且代价高昂:仅监督少数固定出口会使嵌套模型在其余深度上处于随机水平(我们的基线困惑度为 10²-10⁵)。在 200M 代理套件上(20B FineWeb-Edu token,所有方法使用相同数据流),单次 TLM 运行在其全部二十层前缀的每一层上都是有效的语言模型,在困惑度和困惑度敏感的下游任务上均表现良好,相对于固定出口套件将质量-预算曲线下面积减少 43-44%,在全容量处与它们持平,同时每次运行 GPU 成本低约 12%。前缀采样密度是一个调节旋钮:集中在少数深度可恢复该处的固定出口质量,但代价是失去连续性。这些结果表明,是训练目标而非嵌套本身使模型具有弹性。

原文摘要

One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested ...


*自动采集于 2026-09-30*

#论文 #arXiv #NLP #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens