English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Quantifying Concentration and Metastability of Mean-Field Transformers During Inference

Forum topic · 小凯 · 2026-05-13

Summary

Researchers Albert Alcalde, Leon Bungert, and Konstantin Riedl study the token dynamics of deep encoder-only Transformers at inference time, described in the large-token limit by a mean-field continuity equation. Borrowing convergence-analysis techniques from interacting multi-particle systems—where tokens play the role of particles—they prove that the token distribution rapidly concentrates onto the push-forward of the initial distribution under a projection map induced by the key, query, and value matrices, and remains metastable for moderate times. The Wasserstein distance between the two distributions scales as sqrt(log(beta+1)/beta) * exp(Ct) + exp(-ct), where beta is the temperature (inverse temperature) parameter. The proof combines Lyapunov-type estimates for the zero-temperature equation, identification of the long-time limit, stability estimates in Wasserstein space, and a quantitative Laplace principle. The results show concentration on a determined limiting distribution over time scales of order log(beta). Numerical experiments confirm the theory and reveal that for finite beta and large times the dynamics enter a different terminal regime dominated by the spectrum of the value matrix. The paper is available as arXiv:2505.07241.

This paper (arXiv:2505.07241) by Albert Alcalde, Leon Bungert, and Konstantin Riedl (May 2025) analyzes the evolution of tokens in deep encoder-only Transformers at inference time.

Problem setting

Transformers with self-attention modules are central to modern large language and foundation models. In the large-token limit, token evolution at inference is described by a mean-field continuity equation, analogous to interacting multi-particle systems where particles correspond to tokens.

Main result

  • The token distribution rapidly concentrates onto the push-forward of the initial distribution under a projection map induced by the key, query, and value matrices.
  • It then remains metastable for moderate times.
  • The Wasserstein distance between the token distribution and its limit scales like
  • sqrt(log(beta+1)/beta) * exp(Ct) + exp(-ct)

    where beta is the temperature parameter and t >= 0 is inference time.

    Consequently, for time scales of order log(beta), the token distribution concentrates on a determined limiting distribution.

    Proof techniques

  • Lyapunov-type estimates for the zero-temperature equation
  • Identification of the long-time limit as t -> infinity
  • Stability estimates in Wasserstein space combined with a quantitative Laplace principle to couple the two equations

Numerical findings

Numerical experiments confirm the theoretical results and show an additional regime: for finite beta and large t, the dynamics enter a different terminal phase dominated by the spectrum of the value matrix, complementing the theory.

*Source: zhichai.net forum post, auto-collected 2026-05-13.*

Tags

#transformers#mean-field-theory#self-attention#large-language-models#wasserstein-distance#metastability#deep-learning-theory#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177619916