English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Second-Order Path Kernel Interpolation Formulas in Machine Learning

Forum topic · 小凯 · 2026-06-09

Summary

This arXiv paper (2506.08634) by Jin Guo, Roy Y. He, and Jean-Michel Morel extends Pedro Domingos' 2020 path kernel interpolation framework to second order. The original formula expresses a neural network's prediction as an integral, along the optimization path, of a data-dependent kernel aligning model gradients at test and training points. The authors show the leading first-order path-kernel interpolation is supplemented by a curvature-weighted interpolation term. For stochastic gradient descent, an additional sampling-induced component emerges that couples prediction curvature to the covariance of mini-batch gradient noise. The characterization extends to SGD with momentum, where the interpolation structure is preserved but weights are modified by memory-dependent factors. The paper also establishes concentration estimates for final predictions, quantifying fluctuations around the expected second-order representation, refining the path-kernel interpretation of neural network predictions.

Paper Overview

Field: Machine Learning Authors: Jin Guo, Roy Y. He, Jean-Michel Morel Published: 2025-06-11 arXiv: 2506.08634

Abstract (translated from the Chinese forum summary)

Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gradient descent. It expresses the model's prediction as an integral, along the optimization path, of a data-dependent kernel that aligns the model's gradients at the test and training data. Such a first-order characterization remains valid for models trained with batch-based stochastic optimization. This paper develops second-order forms of these interpolation formulas. The authors show that the leading path-kernel interpolation is supplemented by a curvature-weighted interpolation term. For stochastic gradient descent, an additional sampling-induced component appears, coupling prediction curvature with the covariance of mini-batch gradient noise. The characterization is also extended to stochastic gradient descent with momentum, where the interpolation structure is preserved but the weights are modified by memory-dependent factors. Furthermore, the paper establishes concentration estimates for terminal predictions, determining the scale of fluctuations around the expected second-order representation. Together, these results refine the path-kernel interpretation of neural network predictions.

Original Abstract

> Understanding how training data shape neural network predictions is a central problem in modern learning theory. In 2020, Pedro Domingos proposed an interpolation formula valid for every model learned by deterministic gradient descent. It expresses the model's prediction as an integral, along the optimization path, of a data-dependent kernel that aligns the model's gradients at the test and training data. Such a first-order characterization remains valid for models trained with batch-based stochastic optimization. In this paper, we develop second-order forms of these interpolation formulas. We show that the leading path-kernel interpolation is supplemented by a curvature-weighted interpolation term. For stochastic gradient descent, an additional sampling-induced component appears, coupling...

*Source link: https://arxiv.org/abs/2506.08634*

*Auto-collected on 2026-06-09*

Tags

#machine-learning#path-kernel#interpolation-formula#sgd#learning-theory#arxiv#neural-networks

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177981006