Paper Overview
Field: Machine Learning Authors: Jin Guo, Roy Y. He, Jean-Michel Morel Published: 2025-06-11 arXiv: 2506.08634
Abstract (translated from the Chinese forum summary)
Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gradient descent. It expresses the model's prediction as an integral, along the optimization path, of a data-dependent kernel that aligns the model's gradients at the test and training data. Such a first-order characterization remains valid for models trained with batch-based stochastic optimization. This paper develops second-order forms of these interpolation formulas. The authors show that the leading path-kernel interpolation is supplemented by a curvature-weighted interpolation term. For stochastic gradient descent, an additional sampling-induced component appears, coupling prediction curvature with the covariance of mini-batch gradient noise. The characterization is also extended to stochastic gradient descent with momentum, where the interpolation structure is preserved but the weights are modified by memory-dependent factors. Furthermore, the paper establishes concentration estimates for terminal predictions, determining the scale of fluctuations around the expected second-order representation. Together, these results refine the path-kernel interpretation of neural network predictions.
Original Abstract
> Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gradient descent. It expresses the model's prediction as an integral, along the optimization path, of a data-dependent kernel that aligns the model's gradients at the test and training data. Such a first-order characterization remains valid for models trained with batch-based stochastic optimization. In this paper, we develop second-order forms of these interpolation formulas. We show that the leading path-kernel interpolation is supplemented by a curvature-weighted interpolation term. For stochastic gradient descent, an additional sampling-induced component appears, coupling...
*Source link: https://arxiv.org/abs/2506.08634*
*Auto-collected on 2026-06-09*