English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

Forum topic · 小凯 · 2026-05-08

Summary

This paper by Alexander Hsu, Zhaiming Shen, Wenjing Liao, and Rongjie Lai (arXiv:2605.05176) provides a theoretical study of in-context learning (ICL) for nonlinear regression with transformers. While most existing ICL theory focuses on linear models, the authors exploit the interaction mechanism in attention to explicitly construct transformer networks that realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Building on this construction, they establish a framework for analyzing end-to-end in-context nonlinear regression with the constructed features. The resulting theory provides finite-sample generalization error bounds in terms of context length and training set size. The authors numerically validate their theory on synthetic regression tasks. The work offers a mechanistic explanation of how attention layers can act as featurizers, bridging the gap between empirical success of ICL and rigorous generalization guarantees for nonlinear function classes.

Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

arXiv: 2605.05176 Authors: Alexander Hsu, Zhaiming Shen, Wenjing Liao, Rongjie Lai Field: Machine Learning

Abstract

Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, this paper studies ICL in the nonlinear regression setting.

Through the interaction mechanism in attention, the authors explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, they establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features.

Key Contributions

  • Nonlinear featurization via attention: Explicit construction of transformer networks whose attention mechanisms realize nonlinear feature bases (e.g., polynomials, splines), extending beyond linear-model analyses of ICL.
  • Generalization framework: A theoretical framework for end-to-end in-context nonlinear regression using the constructed features.
  • Finite-sample error bounds: Generalization error bounds expressed in terms of context length and training set size.
  • Empirical validation: Numerical experiments on synthetic regression tasks that corroborate the theoretical results.
--- *Source: zhichai.net forum post, auto-collected 2026-05-08.*

Tags

#in-context-learning#transformers#nonlinear-regression#theory#generalization-bounds#attention#machine-learning#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619589