[论文] TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fi...
研究领域: ML 作者: Jichao Jiang, Cristian McGee, El Houcine Bergou, Hanqin Cai, Aritra Dutta 发布时间: 2026-10-01 arXiv: 2610.02199
论文概要
研究领域: ML 作者: Jichao Jiang, Cristian McGee, El Houcine Bergou, Hanqin Cai, Aritra Dutta 发布时间: 2026-10-01 arXiv: 2610.02199
中文摘要
大语言模型(LLM)全参数微调会带来巨大的优化器状态内存开销,限制了现代 GPU 可容纳的模型规模。现有方法要么压缩优化器状态、放弃一阶梯度信息,要么在保留稠密状态的同时改变更新几何结构。近期提出的 Muon 优化器通过矩阵值更新降低优化器内存,但其几何特性与 AdamW 不同,微调 AdamW 预训练模型时可能导致性能下降。为了在 LLM 微调中降低优化器内存且不牺牲精度与计算效率,我们提出三值绝对值最大列稀疏优化器(TACO)。TACO 沿用 Muon 的算子范数最速下降视角并将几何路线推向极致:在维度归一化的 1→1 算子范数下计算精确的最速下降方向——对二维权重矩阵每一列选取绝对值最大元素的符号。这既保留了一阶梯度信息,又使优化器状态内存几乎可忽略。实用的 TACO 优化器每列仅维护少量低精度梯度分量,在 OPT-13B 上相比 AdamW8bit 将持久化优化器状态减少 174 倍(27.7 GB → 0.16 GB),峰值训练内存降低 2.9 倍(80.6 GB → 27.5 GB),同时保持相当的精度与运行速度。TACO 还让 30-32B 参数模型的全参数微调在单张 80 GB H100 GPU 上成为可能,覆盖多个模型家族与任务。
原文摘要
Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO ...
*自动采集于 2026-10-04*
#论文 #arXiv #ML #小凯