Paper Overview
- Field: Computer Vision
- Author: Przemyslaw Musialski
- Published: 2025-06-20
- arXiv: 2506.16802
- The relative geometry of a pair is canonical, \(g_i^{-1} g_j\), so the pairwise invariant \(w_{ij} = \log(g_i^{-1} g_j)\) is intrinsic rather than designed.
- Equivariance under the diagonal \(G\)-action is automatic, and the cocycle condition holds automatically.
- The attention score is the negative squared algebra norm: a canonical proximity kernel under a block-weighted Frobenius inner product, with no need for irreducible representations, spherical harmonics, Clebsch–Gordan products, or learned kernels.
- The closed-form score matches a learned MLP kernel on the same invariant, and outperforms it on SE(2), using 50–80x fewer parameters.
- Vector-token baselines break equivariance by five to twelve orders of magnitude.
Abstract
We place the attention token on the group: a token is an element \(g_i\) of a matrix Lie group \(G\) — a bare transformation, with no feature payload and no external action \(\rho(g)\) carrying it. To our knowledge this is the first attention construction whose tokens are bare matrix Lie group elements: their score is the closed-form algebra norm of the relative pose rather than a learned kernel, and it reaches the affine full-frame groups that every irrep- or surjective-exp-based method must exclude. The construction is called Lie-Algebra Attention.
Once tokens are group elements, the rest follows with none of the usual representation-theoretic machinery:
The construction applies to any matrix Lie group over the chosen log chart, including the non-compact, non-abelian affine group with scale and shear — something no vector-token attention method can reach.
Experimental Results
Three sequence completion experiments on SE(2), SO(3), and Aff(2) confirm the approach:
*Auto-collected on 2026-06-20.*