[论文] TACO: Ternary Absolute-max Column-wise One-sparse Optimizer for LLM Fi...

研究领域: ML 作者: Jichao Jiang, Cristian McGee, El Houcine Bergou, Hanqin Cai, Aritra Dutta 发布时间: 2026-10-01 arXiv: 2610.02199

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: ML 作者: Jichao Jiang, Cristian McGee, El Houcine Bergou, Hanqin Cai, Aritra Dutta 发布时间: 2026-10-01 arXiv: 2610.02199

中文摘要

大语言模型(LLM)全参数微调会带来巨大的优化器状态内存开销,限制了现代 GPU 可容纳的模型规模。现有方法要么压缩优化器状态、放弃一阶梯度信息,要么在保留稠密状态的同时改变更新几何结构。近期提出的 Muon 优化器通过矩阵值更新降低优化器内存,但其几何特性与 AdamW 不同,微调 AdamW 预训练模型时可能导致性能下降。为了在 LLM 微调中降低优化器内存且不牺牲精度与计算效率,我们提出三值绝对值最大列稀疏优化器(TACO)。TACO 沿用 Muon 的算子范数最速下降视角并将几何路线推向极致:在维度归一化的 1→1 算子范数下计算精确的最速下降方向——对二维权重矩阵每一列选取绝对值最大元素的符号。这既保留了一阶梯度信息,又使优化器状态内存几乎可忽略。实用的 TACO 优化器每列仅维护少量低精度梯度分量,在 OPT-13B 上相比 AdamW8bit 将持久化优化器状态减少 174 倍(27.7 GB → 0.16 GB),峰值训练内存降低 2.9 倍(80.6 GB → 27.5 GB),同时保持相当的精度与运行速度。TACO 还让 30-32B 参数模型的全参数微调在单张 80 GB H100 GPU 上成为可能,覆盖多个模型家族与任务。

原文摘要

Full-parameter fine-tuning of large language models (LLMs) incurs substantial optimizer state memory overhead, limiting the model sizes that fit on modern GPUs. Existing approaches either compress optimizer state, abandon first-order gradients, or change the update geometry while retaining dense state. The recently introduced Muon optimizer reduces optimizer memory through matrix-valued updates. Still, its geometry differs from AdamW and can lead to performance degradation when fine-tuning AdamW-pretrained models. To reduce optimizer memory without sacrificing accuracy or computational efficiency in LLM fine-tuning, we propose Ternary Absolute-max Column-wise One-sparse optimizer, or TACO, which follows Muon's operator-norm steepest-descent view but takes the geometric route further. TACO ...


*自动采集于 2026-10-04*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens