The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks
Authors: Shangwen Sun, Alfredo Canziani, Yann LeCun, Jiachen Zhu
- arXiv: 2603.05498
- PDF: https://arxiv.org/pdf/2603.05498.pdf
- Categories: cs.AI, cs.CL
- Field: Natural Language Processing (NLP)
- Type: Empirical study
- Methods: Transformer, Attention
- Massive activations: a small number of tokens exhibit extreme outliers in a few channels.
- Attention sinks: certain tokens attract disproportionate attention mass regardless of semantic relevance.
- Prior work observed that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remained unclear.
- Through systematic experiments, the authors show that the co-occurrence is largely an architectural artifact of modern Transformer design.
- The two phenomena serve related but distinct functions.
- Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit parameters.
Overview
This paper studies two recurring phenomena in Transformer language models:
Key Findings
Impact Assessment
The study carries significant theoretical and practical value and may notably influence related research areas.
Abstract (excerpts from original)
> We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, in which certain tokens attract disproportionate attention mass regardless of semantic relevance. Prior work observes that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remain unclear. Through systematic experiments, we show that the co-occurrence is largely an architectural artifact of modern Transformer design, and that the two phenomena serve related but distinct functions. Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit paramete...
*Auto-collected on 2026-03-07.*