Paper Overview
Field: Computer Vision Author: Przemyslaw Musialski Published: 2025-06-23 arXiv: 2506.18493
Summary
This paper puts the attention token on the group: a token is an element g_i of a matrix Lie group G — a bare transformation, with no feature payload and no external action rho(g) carrying it. According to the author, this is the first attention construction whose tokens are bare matrix Lie group elements: their score is the closed-form algebra norm of the relative pose rather than a learned kernel, and it reaches the affine full-frame groups that every irrep- or surjective-exp-based method must exclude. The construction is called Lie-Algebra Attention.
Once tokens are group elements, the rest follows without any of the usual representation-theoretic machinery:
- The relative geometry of a pair is canonical, g_i^{-1} g_j, so the pairwise invariant w_{ij} = log(g_i^{-1} g_j) is intrinsic rather than designed.
- Equivariance under the diagonal G-action is tautological, and the cocycle condition holds automatically.
- The attention score is the negative squared algebra norm: s_{ij} = -||log(g_i^{-1} g_j)||_λ^2 / τ — a canonical proximity kernel under a block-weighted Frobenius inner product, with no irreducible representations, spherical harmonics, Clebsch-Gordan products, or learned kernels.
- The closed-form scores match learned MLP kernels on the same invariants, and outperform them on SE(2), using 50 to 80 times fewer score parameters.
- Vector-token baselines break invariance by five to twelve orders of magnitude.
The construction applies to any matrix Lie group whose logarithm contains the relative pose, including non-compact non-abelian affine groups (with scale and shear), which no vector-token attention method can reach.
Experimental Results
Three sequence completion experiments on SE(2), SO(3), and Aff(2) confirm the claims:
Original Abstract (excerpt)
> We place the attention token on the group: a token is an element g_i of a matrix Lie group G — a bare transformation, with no feature payload and no external action rho(g) carrying it. To our knowledge this is the first attention construction whose tokens are bare matrix Lie group elements: their score is the closed-form algebra norm of the relative pose rather than a learned kernel, and it reaches the affine full-frame groups that every irrep- or surjective-exp-based method must exclude. We call it Lie-Algebra Attention. Once tokens are group elements, the rest follows with none of the usual representation-theoretic machinery. The relative geometry of a pair is canonical, g_i^{-1} g_j, so the pairwise invariant w_{ij} = log(g_i^{-1} g_j) is intrinsic rather than designed; equivariance un...
--- *Auto-collected on 2026-06-23*