Attention Residuals Explained: When Residual Connections Meet Attention Mechanisms
In March 2026, the Kimi team released a technical report titled Attention Residuals, challenging the foundational residual connection that has defined deep neural networks since ResNet (2015). The paper's first author, Guangyu Chen, is a 17-year-old high school senior from Shenzhen who interned at Kimi for five months. The release drew public reactions from Elon Musk ("Impressive work from Kimi"), Andrej Karpathy ("We still haven't taken 'Attention is All You Need' literally enough"), and former OpenAI researcher Jerry Tworek ("Everything needs to be rethought, Deep Learning 2.0 is coming").
Key points
- The problem — "PreNorm Dilution": Traditional residual connections equally accumulate every layer's output (
h_l = h_{l-1} + f_l(Norm(h_{l-1}))). As depth grows, hidden-state magnitude keeps rising, early-layer information gets drowned out, and later layers need increasingly strong signals to matter — analogous to working-memory overload. - The core idea: Apply attention over *depth* instead of (or in addition to) *sequence*. Each layer computes
h_l = Σ_i α_{l→i} · v_i, where softmax attention weights dynamically decide which historical layers to aggregate from — a duality with how Transformers attend over tokens. - Two implementations:
- Full AttnRes: Each layer maintains a pseudo-query vector and attends over all previous layers' outputs. Maximum expressivity, but O(L·d) memory and heavy cross-server communication.
- Block AttnRes: Standard residuals within blocks; attention residuals across block summaries. Memory drops to O(N·d), with a two-phase inference scheme (parallel inter-block attention, then sequential intra-block attention with online softmax merge) keeping inference latency overhead under 2%.
- Scaling law: Block AttnRes reaches the same loss with only 80% of the baseline's compute — effectively 1.25x free compute.
- Downstream benchmarks: +7.5% on GPQA-Diamond, +3.6% on math reasoning, +3.1% on HumanEval, plus improvements on MMLU. Gains are strongest on multi-step reasoning tasks.
- Internal analysis: Output magnitudes stay bounded across depth, gradients distribute more evenly across layers, and attention maps show both local and long-range jump connections with functional specialization of layers.
- Paper: https://arxiv.org/abs/2603.15031
- Official code: https://github.com/MoonshotAI/Attention-Residuals
- Community implementation (kyegomez/attn_res): https://github.com/kyegomez/attn_res
- Related: Kimi Linear (arXiv:2510.26692)
Experimental results
Experiments used Kimi Linear (48B total / 3B active parameters) trained on 1.4T tokens:
Relation to prior work
| Method | Mechanism | Dynamism | |--------|-----------|----------| | DenseFormer (2024) | Learnable static scalar weights | Static | | Hyper-Connections (2025, ByteDance Seed) | Multi-stream residual expansion | Partially dynamic | | mHC (2025, DeepSeek) | Birkhoff manifold-constrained mixing | Geometric constraint | | AttnRes (2026) | Softmax attention | Fully dynamic |
Why it matters
Attention Residuals turns network connectivity from static to dynamic and input-adaptive, potentially removing the practical depth ceiling imposed by PreNorm dilution. Where the 2017 Transformer brought attention to the sequence dimension, AttnRes extends it to the depth dimension — a possible first step toward what the community is calling "Deep Learning 2.0."