English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

When Attention Goes Blind: A Numerical Underflow Bug Hidden in ALiBi Positional Encoding

Forum topic · ✨步子哥 · 2026-08-05

Summary

A 2026 arXiv paper, 'When Attention Goes Blind,' reveals that ALiBi positional encoding suffers from floating-point underflow: when token distance grows large, the linear attention bias pushes softmax inputs below the underflow threshold (approximately -103.27 for fp32 and -92.18 for bf16), zeroing attention weights and making attention heads unable to see distant tokens. Heads with steeper slopes go blind first, so models fail progressively rather than all at once. Controlled experiments on a 148M-parameter decoder show standard metrics like perplexity remain unaffected, while passkey retrieval degrades severely; notably, default ALiBi slopes still perform well on needle-in-a-haystack tests. Among four mitigation strategies tested (clamping, robust slopes, log-scaled distances, soft capping), log-scaled distances most consistently improve passkey retrieval. The finding illustrates how single-metric evaluation hides critical failure modes and warns that published ALiBi-based models may already be affected at long ranges.

When Attention Goes Blind: A Numerical Underflow Bug Hidden in ALiBi Positional Encoding

You trained a language model with ALiBi positional encoding. Context window 8K, training loss normal, standard benchmark scores normal. You think everything is fine.

But your model may already be partially blind on needle-in-a-haystack retrieval tasks — some attention heads simply cannot see distant tokens at long ranges, because floating-point underflow zeroes out their attention weights.

This is not hypothetical. It is the core finding of an August 2026 arXiv paper, "When Attention Goes Blind": ALiBi's linear bias underflows in floating-point precision, turning a swath of attention weights into zeros, so attention heads can no longer "see" faraway tokens.

What ALiBi Is, and Why It Goes Blind

A quick recap of ALiBi. In Transformer attention, each token attends to all other tokens. The original Transformer used sinusoidal positional encoding; RoPE later became mainstream. ALiBi, proposed by Press et al. in 2022, is an alternative: instead of adding positional encodings to token embeddings, it adds a linear bias to attention scores — the farther apart tokens are, the larger (negative) the bias, and the smaller the attention weight.

Intuitively reasonable: distant tokens matter less than nearby ones, and linear decay is a simple inductive bias.

The problem lies in the numerical implementation of softmax. Attention weights are computed as:

\[\text{softmax}(A + B)\]

where \(A\) is the attention score and \(B = -m \cdot D\) is the bias (\(m\) is the per-head slope, \(D\) is the distance matrix). When distance \(D\) is large enough, \(B\) becomes a large negative number, pushing \(A + B\) very negative. Softmax takes \(e^x\) of every element — when \(x\) is sufficiently negative, \(e^x\) underflows to zero.

That is the blindness: beyond a distance threshold, distant tokens' exponentials in softmax become exactly zero, their attention weights are wiped out, and the head cannot see them anymore.

When It Triggers

The paper gives concrete thresholds:

  • fp32 (single precision): underflow threshold \(\tau_{u,\text{fp32}} \approx -103.27\)
  • bf16 (bfloat16): underflow threshold \(\tau_{u,\text{bf16}} \approx -92.18\)
  • Different ALiBi heads use different slopes. The steeper the slope, the faster the decay and the earlier the underflow. Per the paper: "the steepest-slope head underflows first."

    This means in a multi-head model, not all heads go blind simultaneously — they go blind one by one, from steepest to gentlest slope. Some heads can no longer see faraway tokens before the context even reaches 8K, while others hold out longer. The result: the model appears to work, but some heads have become "nearsighted."

    How Big Is the Impact

    The paper ran controlled experiments with a 148M-parameter decoder, separating this failure mode from "out-of-context degradation" (another known issue). Conclusions:

    1. Little impact on standard benchmarks. Perplexity and downstream task scores barely move. That is why this bug went unnoticed: standard benchmarks can't see it.

    2. Big impact on token retrieval. On passkey retrieval (finding a password in a long text), ALiBi blindness severely hurts performance. Makes sense — retrieval tasks specifically need heads to see distant tokens, and blind heads can't.

    3. But on needle-in-a-haystack, default ALiBi slopes remain a strong baseline. A counterintuitive finding: despite the underflow, the default slope configuration still performs well on NIAH retrieval. The paper's explanation: the default slopes happen to place underflow at "not-so-critical" distances.

    Four Fixes

    The paper tested four training-time mitigation strategies:

    1. Clamping: cap the bias at a maximum value so it cannot grow indefinitely 2. Robust Slopes: redesign the slope sequence to avoid overly fast decay 3. Log-scaled Distances: replace the linear bias with logarithmic decay — slower falloff at long range, avoiding unbounded negative drift 4. Soft Capping: replace hard clamping with a smooth function

    Result: log-scaled distances most consistently improved passkey retrieval performance.

    The intuition behind log decay: a linear bias assumes "importance halves every time distance doubles" — too aggressive. A log bias assumes "importance halves every order of magnitude of distance" — closer to how relevance actually decays in real scenarios.

    What This Tells Us

    This finding is another instance of the "evaluation blind-spot law": failure modes invisible to standard benchmarks are exactly where bugs hide.

    ALiBi blindness doesn't affect perplexity or standard downstream scores, so standard evaluation says "all clear." But test long-range retrieval specifically, and the problem surfaces. This is the same class of problem as Progressive Cramming (99% token accuracy masking 100% generation failure) and TriviaRoomQA (fine inside the knowledge boundary, random beyond it): a single metric can mask critical failure modes.

    The deeper cause: ALiBi's linear bias assumes unbounded decay — distance can grow forever, the bias can grow forever negative. But floating point has precision limits. When the math assumes "infinite" and the implementation is "finite," a mismatch appears at the boundary. This is the same class of issue as RoPE's precision problems under bfloat16 (Wang et al. 2025): a crack between the mathematical assumptions of positional encodings and the physical reality of floating-point arithmetic.

    Practical Advice

    The paper offers concrete training recommendations:

  • If you use ALiBi, consider switching to log-distance bias
  • If you train new models, monitor passkey retrieval, not just perplexity
  • If you deploy existing ALiBi models for long-context applications, be aware some heads may already be blind at your operating distances
  • Default slope configurations are fine on NIAH, but check passkey retrieval separately
The paper notes this problem appears in "ALiBi-based SOTA pretrained models" — it's not only an issue when training new models; already-released models may be affected too.

An Engineering Lesson

This bug stayed hidden so long because it appears only on specific tasks, at specific distances, in a specific way. The "average score" of standard benchmarks dilutes it.

For everyone designing new positional encodings, the lesson is: your positional encoding working at 1K context does not mean it works at 32K context. Floating-point precision is not infinite, and your formulas may diverge from the implementation at the boundaries.

Test long-range retrieval, not just standard benchmarks. Your model may be more "nearsighted" than you think.

---

Paper: https://arxiv.org/abs/2608.03994

Code: The paper states "Code will be released upon publication"

Tags

#alibi#positional-encoding#attention#numerical-underflow#long-context#needle-in-a-haystack#transformers#llm

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178595033