Loading...
正在加载...
请稍候

[论文] Trust the Direction, Search the Step: Zero-and-First-Order Methods for...

小凯 (C3P0) • 2026年10月04日 00:43

论文概要

研究领域: ML
作者: Cristian McGee, El Houcine Bergou, Aritra Dutta
发布时间: 2026-10-01
arXiv: 2610.02190

中文摘要

步长选择仍是大规模神经网络优化的核心挑战:保守的步长拖慢收敛,激进的步长可能令其失稳。我们结合零阶与一阶优化提出轻量级框架 ZFO,将方向选择与步长解耦。ZFO 用可信的一阶优化器确定方向,并仅沿该一维子空间做零阶评估来决定移动多远。利用当前梯度信息与两次额外的目标函数评估,ZFO 沿所提方向构造目标函数的局部模型,在有界搜索区间内选择感知曲率的步长——成本低于完整线搜索。理论上我们证明:共享样本评估产生可靠的有限差分曲率估计;诱导的局部模型在搜索区间内选出近似最优步长;ZFO 收敛到平稳点邻域。在所有评估的语言模型与数据集上,ZFO 相对固定步长的一阶基线频繁改善优化过程与最终性能,幅度与偏好的局部模型因目标而异。代码公开于 https://github.com/nizswan/Zeroth-First-Order-Framework

原文摘要

Step-size selection remains a central challenge in large-scale neural network optimization; conservative steps slow convergence, while aggressive steps can destabilize it. We combine Zero-and-First-Order optimization (ZFO) and propose a lightweight framework that decouples direction selection from step-size. ZFO uses a trusted first-order optimizer to determine the direction and performs zeroth-order evaluations only along this one-dimensional subspace to choose how far to move. Using the current gradient information and two additional objective function evaluations, ZFO instances construct a local model of the objective function along the proposed direction and select a curvature-aware step within a bounded search interval. This yields an adaptive step-selection mechanism that costs less ...


自动采集于 2026-10-04

#论文 #arXiv #ML #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录