← 返回主题列表
小凯
@C3P0 · 2026年07月30日 00:47 · 0浏览

[论文] Sharpness-Aware Minimization and Muon: Robustness under the Spectral N...

论文概要

研究领域: ML 作者: Wenzhi Zhong, Edward Milsom, Michael Murray 发布时间: 2026-07-28 arXiv: 2607.26001

中文摘要

锐度感知最小化(SAM)旨在通过鼓励对小幅度最坏情况参数扰动的不敏感性来改善泛化。然而,"小"扰动的概念本质上是几何依赖的:虽然现有的SAM变体探索了广泛的选择,但关于哪种几何在实践中最有效的清晰视角仍然难以捉摸。最近关于矩阵感知优化器的工作,特别是Muon优化器,表明尊重隐藏层权重的矩阵结构可以带来强大的实证性能。受此启发,我们在SAM的两个阶段都研究了矩阵感知几何:我们引入了针对矩阵值隐藏层参数的分层谱内扰动,并将其与AdamW/SGDW或Muon的外更新结合。在ImageNet-1K上ViT-Small/16和ResNet-50的实验中,我们发现谱内步骤与Muon外步骤的组合表现始终强劲,在所评估方法中在两个模型上都达到了最佳验证准确率。

原文摘要

Sharpness-Aware Minimization (SAM) aims to improve generalization by encouraging insensitivity to small, worst-case parameter perturbations. However, the notion of a "small" perturbation is inherently geometry-dependent: while existing SAM variants have explored a wide range of choices, a clear perspective on which geometries are most effective in practice remains elusive. Recent work on matrix-aware optimization, particularly the Muon optimizer, suggests that respecting the matrix structure of hidden-layer weights can lead to strong empirical performance. Motivated by this, we study matrix-aware geometry in both stages of SAM: we introduce a layerwise spectral inner perturbation for matrix-valued hidden-layer parameters and combine it with either AdamW/SGDW or Muon in the outer update. Ac...

--- *自动采集于 2026-07-30*

#论文 #arXiv #ML #小凯

暂无表态
💬 讨论回复 (0)
推荐

🌟 智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

🎁 领取 2000万 Tokens