English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Understanding In-Context Learning for Nonlinear Regression with Transformers: Attention as Featurizer

Forum topic · 小凯 · 2026-05-08

Summary

This post introduces an arXiv paper (2605.05176) by Alexander Hsu, Zhaiming Shen, Wenjing Liao, and Rongjie Lai on the theory of in-context learning (ICL) for nonlinear regression. While most existing theoretical work on ICL focuses on linear models, the authors study how pre-trained transformers can learn from examples in the prompt without weight updates in nonlinear settings. They explicitly construct transformer networks whose attention mechanism realizes nonlinear features, such as polynomial or spline bases, spanning a broad class of functions. Building on this construction, they establish a framework for analyzing end-to-end in-context nonlinear regression with the constructed features, deriving finite-sample generalization error bounds in terms of context length and training set size. The theory is numerically validated on synthetic regression tasks.

Paper Overview

Field: Machine Learning Authors: Alexander Hsu, Zhaiming Shen, Wenjing Liao, Rongjie Lai Published: 2026-05-06 arXiv: 2605.05176

Abstract

Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory has focused on linear models, this paper studies ICL in the nonlinear regression setting.

Through the interaction mechanism in attention, the authors explicitly construct transformer networks to realize nonlinear features, such as polynomial or spline bases, which span a wide class of functions. Based on this construction, they establish a framework to analyze end-to-end in-context nonlinear regression with the constructed features.

Key Contributions

  • Explicit construction of transformer networks that realize nonlinear feature bases (polynomial, spline) via attention interactions
  • A theoretical framework for analyzing end-to-end in-context nonlinear regression with these constructed features
  • Finite-sample generalization error bounds in terms of context length and training set size
  • Numerical validation of the theory on synthetic regression tasks

Link

Full paper: https://arxiv.org/abs/2605.05176

--- *Auto-collected on 2026-05-08*

Tags

#in-context-learning#transformers#nonlinear-regression#machine-learning#theory#arxiv#attention-mechanism

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619589