论文概要
研究领域: ML
作者: Anish Kataria
发布时间: 2026-09-11
arXiv: 2509.05824
中文摘要
过拟合记忆后训练的神经网络常经历向泛化的延迟过渡,这一现象称为grokking。尽管对其发生原因已有理论进展,但该过渡在超参数空间中何时发生的定量结构仍未被刻画。我们在模运算上针对384种双层MLP配置绘制记忆-泛化边界,拟合泛化 onset 时间的幂律缩放关系:\(T_{\mathrm{grok}} \propto H^{-0.27} D^{-2.04} \eta^{-0.50} \lambda^{-0.64}\) (\(R^2=0.732\);含交互项$0.821$)。指数层级揭示数据复杂度(\(D^{-2.04}\))是状态转换的主导驱动,而非模型容量(\(H^{-0.27}\)):数据翻倍加速泛化约4倍,而宽度翻倍仅约1.2倍。权重衰减 \(\lambda \gtrsim 1.0\) 处存在尖锐相边界,分隔grokking与非grokking配置,权重范数轨迹显示过渡期间单调压缩,与隐式正则化选择低复杂度解一致。这些结果为预测和控制过参数化网络中的状态转换提供定量基础。
原文摘要
Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: \(T_{\mathrm{grok}} \propto H^{-0.27}\, D^{-2.04}\, \eta^{-0.50}\, \lambda^{-0.64}\) (\(R^2 = 0.732\); $0.821$ with interactions). The exponent hierarchy reveals that data复杂度 (\(D^{-2.04}\)) is the dominant driver of regime transition, not model capacity (\(H^{-0.27}\)): doubling data accelerates generalization b...
自动采集于 2026-09-12
#论文 #arXiv #ML #小凯
讨论回复
加载中...正在加载回复...
推荐
智谱 GLM-5 已上线
我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。