Paper Overview
Field: cs.LG, cs.CV, cs.GR Author: Przemyslaw Musialski Published: 2026-06-21 arXiv: 2506.17582
Summary
This paper introduces Lie-algebra attention, where each attention token is an element \(g_i\) of a matrix Lie group \(G\) — a bare transformation carrying no feature payload and no external representation action \(\rho(g)\). To the authors' knowledge, this is the first attention construction that uses bare matrix Lie group elements as tokens.
Key points
- Tokens as group elements: Each token is a raw transformation in a matrix Lie group \(G\), with no feature load and no external action \(\rho(g)\).
- Canonical pairwise geometry: Relative poses are given by \(g_i^{-1} g_j\), so the pairwise invariant \(w_{ij} = \log(g_i^{-1} g_j)\) is intrinsic rather than designed. Equivariance under the diagonal \(G\) action is automatic, and the cocycle condition is naturally satisfied.
- Closed-form attention scores: Scores take the form
- Broad applicability: The construction works for any matrix Lie group, provided the chosen log chart covers relative poses — including the non-compact, non-abelian affine group, which contains scale and shear and is unreachable by any vector-token attention method, whether based on irreducible representations or surjective exponential maps.
- Experimental validation: Three sequence completion experiments on SE(2), SO(3), and Aff(2) show that the closed-form scores match learned MLP kernels on the same invariants, outperform them on SE(2) with 50–80× fewer score parameters, while vector-token baselines break equivariance by five to twelve orders of magnitude.
a canonical proximity kernel under a block-weighted Frobenius inner product — no irreducible representations, spherical harmonics, Clebsch–Gordan products, or learned kernels required.
*Auto-collected on 2026-06-21*