Quantifying the Memorization-to-Generalization Transition: Scaling Laws and Phase Structure in Grokking

研究领域: ML 作者: Anish Kataria 发布时间: 2026-09-11 arXiv: 2509.05824

论文概要

研究领域: ML 作者: Anish Kataria 发布时间: 2026-09-11 arXiv: 2509.05824

中文摘要

过拟合记忆后训练的神经网络常经历向泛化的延迟过渡,这一现象称为grokking。尽管对其发生原因已有理论进展,但该过渡在超参数空间中何时发生的定量结构仍未被刻画。我们在模运算上针对384种双层MLP配置绘制记忆-泛化边界,拟合泛化 onset 时间的幂律缩放关系:\(T_{\mathrm{grok}} \propto H^{-0.27} D^{-2.04} \eta^{-0.50} \lambda^{-0.64}\) (\(R^2=0.732\);含交互项$0.821$)。指数层级揭示数据复杂度(\(D^{-2.04}\))是状态转换的主导驱动,而非模型容量(\(H^{-0.27}\)):数据翻倍加速泛化约4倍,而宽度翻倍仅约1.2倍。权重衰减 \(\lambda \gtrsim 1.0\) 处存在尖锐相边界,分隔grokking与非grokking配置,权重范数轨迹显示过渡期间单调压缩,与隐式正则化选择低复杂度解一致。这些结果为预测和控制过参数化网络中的状态转换提供定量基础。

原文摘要

Neural networks trained past memorization frequently undergo a delayed transition to generalization, a phenomenon known as grokking. Despite theoretical progress on \emph{why} this transition occurs, the quantitative structure of \emph{when} it occurs in hyperparameter space remains uncharacterized. We map the memorization-to-generalization boundary across 384 configurations of two-hidden-layer MLPs on modular arithmetic, fitting a power-law scaling relation for generalization onset time: \(T_{\mathrm{grok}} \propto H^{-0.27}\, D^{-2.04}\, \eta^{-0.50}\, \lambda^{-0.64}\) (\(R^2 = 0.732\); $0.821$ with interactions). The exponent hierarchy reveals that data复杂度 (\(D^{-2.04}\)) is the dominant driver of regime transition, not model capacity (\(H^{-0.27}\)): doubling data accelerates generalization b...


*自动采集于 2026-09-12*

#论文 #arXiv #ML #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens