论文概要
研究领域: NLP
作者: Shouren Wang, Chuang Ma, Mohsen Hariri, Debargha Ganguly, Wang Yang, Xiaoqing Tong, Qianying Liu, Xiaotian Han, Vipin Chaudhary
发布时间: 2026-09-28
arXiv: 2609.35751
中文摘要
循环 Transformer 多次复用一块层:通过花费额外计算,将固定大小的模型推向更远处,从而更充分地利用其参数;而稀疏 MoE 模型为每个 token 仅激活众多专家中的少数几个。循环 MoE 桥接了这两种设计理念,为 MoE 模型提供了更好利用专家的新潜力,但这也引出一个问题:如何循环一个 MoE?我们用 Foil 回答:在专家参数和每 token 专家计算量固定的情况下,Foil (1) 扁平化专家——将专家层数减半、每层专家数加倍、通过次数加倍,使每次路由决策从更大的池中选择;(2) 解绑注意力——为每次通过赋予独立的注意力参数,而专家和路由器保持共享。实验表明 Foil 明显优于未扁平化的循环基线:在 20B token 时每个 Foil 模型的预训练损失都更低;在 100B token 时损失随扁平化程度单调改善,最扁平的 Foil 在同等参数和计算下比基线低 0.012 nat,下游准确率持平或更好;解绑注意力还在同等形状下产生更均衡、更自信的路由。我们的消融分析了 Foil 为何有效,并将其转化为循环 MoE 的设计指导:循环和加宽专家层的收益相互放大,路由信心比负载均衡更能追踪健康的专家使用,因此稀疏循环 MoE 应该使用更多每层专家和更多通过次数。
原文摘要
Looped Transformers reuse one block of layers several times: by spending extra computation they push a model of fixed size further, and so use its parameters more fully; while sparse mixture-of-experts (MoE) models activate only a few of many experts for each token. Looped MoE bridges these two design philosophies and gives MoE models new potential for better expert usage, but it raises a question: how to loop a MoE? We answer it with Foil. With the expert parameters and the expert compute per token held fixed, Foil (1) flattens the experts, halving the expert layers, doubling the experts per layer and doubling the passes, so that every routing decision chooses from a larger pool, and (2) unties the attention, giving each pass its own attention parameters while the experts and routers stay...
自动采集于 2026-09-30
#论文 #arXiv #NLP #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。