English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks in Transformers

Forum topic · 小凯 · 2026-03-07

Summary

This arXiv paper (2603.05498) by Shangwen Sun, Alfredo Canziani, Yann LeCun, and Jiachen Zhu investigates two recurring phenomena in Transformer language models: massive activations, where a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, where certain tokens attract disproportionate attention mass regardless of semantic relevance. Although prior work observes that these phenomena frequently co-occur and often involve the same tokens, their functional roles and causal relationship remained unclear. Through systematic experiments, the authors show that the co-occurrence is largely an architectural artifact of modern Transformer design, and that the two phenomena serve related but distinct functions. Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit parameters. The study provides important theoretical and practical insights for understanding Transformer architecture and NLP research.

The Spike, the Sparse and the Sink: Anatomy of Massive Activations and Attention Sinks

Authors: Shangwen Sun, Alfredo Canziani, Yann LeCun, Jiachen Zhu

  • arXiv: 2603.05498
  • PDF: https://arxiv.org/pdf/2603.05498.pdf
  • Categories: cs.AI, cs.CL
  • Field: Natural Language Processing (NLP)
  • Type: Empirical study
  • Methods: Transformer, Attention
  • Overview

    This paper studies two recurring phenomena in Transformer language models:

  • Massive activations: a small number of tokens exhibit extreme outliers in a few channels.
  • Attention sinks: certain tokens attract disproportionate attention mass regardless of semantic relevance.
  • Key Findings

  • Prior work observed that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remained unclear.
  • Through systematic experiments, the authors show that the co-occurrence is largely an architectural artifact of modern Transformer design.
  • The two phenomena serve related but distinct functions.
  • Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit parameters.

Impact Assessment

The study carries significant theoretical and practical value and may notably influence related research areas.

Abstract (excerpts from original)

> We study two recurring phenomena in Transformer language models: massive activations, in which a small number of tokens exhibit extreme outliers in a few channels, and attention sinks, in which certain tokens attract disproportionate attention mass regardless of semantic relevance. Prior work observes that these phenomena frequently co-occur and often involve the same tokens, but their functional roles and causal relationship remain unclear. Through systematic experiments, we show that the co-occurrence is largely an architectural artifact of modern Transformer design, and that the two phenomena serve related but distinct functions. Massive activations operate globally: they induce near-constant hidden representations that persist across layers, effectively functioning as implicit paramete...

*Auto-collected on 2026-03-07.*

Tags

#transformer#attention-sinks#massive-activations#nlp#arxiv#deep-learning#language-models

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177168727