论文概要
研究领域: ML
作者: Zeyan Li, Panqi Yang, Qirong Guo, Shengda Zhuo, SIyuan Qiu, Hu Xu, Chun Li, Jianfeng Xu
发布时间: 2026-09-25
arXiv: 2609.31600
中文摘要
低秩适配器(LoRA)让大语言模型按任务廉价微调成为可能,但把多个独立训练的适配器合并进一个模型仍然困难:在权重空间直接合并更新会引起干扰,在所有任务数据上重训代价高昂,而在多个独立适配器之间做路由则背离了“单一合并模型”的目标。我们将困难追溯到每种组合方法都在隐式做出的两个选择。其一,LoRA 更新存在无穷多个等价分解;当适配器单独使用时这一选择不可见,但它决定了适配器之间学习到的交互“能看到什么”。其二,新旧技能之间的耦合可以指向任一方向,而方向决定了旧技能能否继续计算它原本计算的东西。我们提出 READ(Read-only Expansion of Adapter Deltas,适配器增量的只读扩展),同时固定这两个选择:每个适配器被重写为一种精确保持其更新的平衡典范形式;耦合只单向增长——新技能可以读取旧技能的输入子空间,但不能写入它们的输出子空间。每次追加新技能时唯一可训练的对象是新技能在耦合矩阵中的那一行,且合并后的更新直接折叠进基座权重,无推理开销、无需路由、无需任务特定规则。我们在四个基准套件和两个模型家族上评测 READ,每次只增加一个技能。READ 在多个模型家族上将每个套件的平均分都提升超过由相同适配器构建的最强已发表基线——SuperGLUE 提升逾 20 分、领域套件提升逾 7 分——且几乎所有完整的“逐次叠加”序列最终都高于所有直接基线。单个适配器从不暴露的分解坐标与耦合方向,正是决定组合技能能否存活的关键。
原文摘要
Low-rank adapters (LoRA) make it cheap to fine-tune a large language model once per task, but combining several independently trained adapters into one model remains difficult: merging the updates in weight space causes interference, retraining on all task data is expensive, and routing between separate adapters gives up the goal of a single combined model. We trace the difficulty to two choices that every composition method makes implicitly. A LoRA update admits infinitely many equivalent factorizations; the choice among them is invisible while an adapter serves alone, but it determines what a learned interaction between adapters can see. A coupling between an old skill and a new one can likewise point in either direction, and the direction decides whether the old skills keep computing wh...
自动采集于 2026-09-29
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。