English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Attention Residuals Explained: When Residual Connections Meet Attention Mechanisms

Forum topic · 小凯 · 2026-04-05

Summary

Attention Residuals (AttnRes), a new architecture technique from the Kimi (Moonshot AI) team, replaces the decade-old fixed residual connection in deep neural networks with softmax-based attention across depth. Instead of equally accumulating every layer's output—which the authors identify as 'PreNorm Dilution' that buries early-layer information and inflates hidden-state magnitude—each layer dynamically selects which previous layers to aggregate from, mirroring how Transformers apply attention over sequences. The paper (arXiv:2603.15031) proposes two variants: Full AttnRes, offering maximum expressivity at O(L·d) memory cost, and Block AttnRes, which applies traditional residuals within blocks and attention residuals across block summaries, reducing memory to O(N·d) with under 2% inference latency overhead. Experiments on Kimi Linear (48B total / 3B active parameters, 1.4T tokens) show Block AttnRes matches baseline loss with only 80% of the compute, plus gains of +7.5% on GPQA-Diamond, +3.6% on math reasoning, and +3.1% on HumanEval. The work draws comparisons to DenseFormer, Hyper-Connections, and DeepSeek's mHC, and drew praise from Elon Musk and Andrej Karpathy, with code open-sourced on GitHub.

Attention Residuals Explained: When Residual Connections Meet Attention Mechanisms

In March 2026, the Kimi team released a technical report titled Attention Residuals, challenging the foundational residual connection that has defined deep neural networks since ResNet (2015). The paper's first author, Guangyu Chen, is a 17-year-old high school senior from Shenzhen who interned at Kimi for five months. The release drew public reactions from Elon Musk ("Impressive work from Kimi"), Andrej Karpathy ("We still haven't taken 'Attention is All You Need' literally enough"), and former OpenAI researcher Jerry Tworek ("Everything needs to be rethought, Deep Learning 2.0 is coming").

Key points

  • The problem — "PreNorm Dilution": Traditional residual connections equally accumulate every layer's output (h_l = h_{l-1} + f_l(Norm(h_{l-1}))). As depth grows, hidden-state magnitude keeps rising, early-layer information gets drowned out, and later layers need increasingly strong signals to matter — analogous to working-memory overload.
  • The core idea: Apply attention over *depth* instead of (or in addition to) *sequence*. Each layer computes h_l = Σ_i α_{l→i} · v_i, where softmax attention weights dynamically decide which historical layers to aggregate from — a duality with how Transformers attend over tokens.
  • Two implementations:
  • Full AttnRes: Each layer maintains a pseudo-query vector and attends over all previous layers' outputs. Maximum expressivity, but O(L·d) memory and heavy cross-server communication.
  • Block AttnRes: Standard residuals within blocks; attention residuals across block summaries. Memory drops to O(N·d), with a two-phase inference scheme (parallel inter-block attention, then sequential intra-block attention with online softmax merge) keeping inference latency overhead under 2%.
  • Experimental results

    Experiments used Kimi Linear (48B total / 3B active parameters) trained on 1.4T tokens:

  • Scaling law: Block AttnRes reaches the same loss with only 80% of the baseline's compute — effectively 1.25x free compute.
  • Downstream benchmarks: +7.5% on GPQA-Diamond, +3.6% on math reasoning, +3.1% on HumanEval, plus improvements on MMLU. Gains are strongest on multi-step reasoning tasks.
  • Internal analysis: Output magnitudes stay bounded across depth, gradients distribute more evenly across layers, and attention maps show both local and long-range jump connections with functional specialization of layers.
  • Relation to prior work

    | Method | Mechanism | Dynamism | |--------|-----------|----------| | DenseFormer (2024) | Learnable static scalar weights | Static | | Hyper-Connections (2025, ByteDance Seed) | Multi-stream residual expansion | Partially dynamic | | mHC (2025, DeepSeek) | Birkhoff manifold-constrained mixing | Geometric constraint | | AttnRes (2026) | Softmax attention | Fully dynamic |

    Why it matters

    Attention Residuals turns network connectivity from static to dynamic and input-adaptive, potentially removing the practical depth ceiling imposed by PreNorm dilution. Where the 2017 Transformer brought attention to the sequence dimension, AttnRes extends it to the depth dimension — a possible first step toward what the community is calling "Deep Learning 2.0."

    Resources

  • Paper: https://arxiv.org/abs/2603.15031
  • Official code: https://github.com/MoonshotAI/Attention-Residuals
  • Community implementation (kyegomez/attn_res): https://github.com/kyegomez/attn_res
  • Related: Kimi Linear (arXiv:2510.26692)
*Note: Claims and figures above are as reported in the source forum post and the referenced paper.*

Tags

#attention-residuals#kimi#transformer#residual-connections#deep-learning#model-architecture#scaling-laws#open-source

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169567