An Invisible Blind Spot
Imagine you trained a 7B-parameter language model with a 2048-token context window using ALiBi positional encoding—cheap, parameter-free, and capable of extrapolating to longer contexts. Standard evaluations pass, perplexity is normal, and CS/QA/LG benchmarks differ by only 1.6 to 3.4 percentage points. You relax and ship the model.
Then someone places the sentence "the password is 42-7B" 3000 tokens before the question and asks your model for the password. It fails.
It's not stupid. Its attention heads, starting from token 2048, simply cannot see the earlier password. Not blurred—completely invisible: the corresponding attention weights have been zeroed out by floating-point precision. The model doesn't know it's blind, you don't know, and standard benchmarks can't detect it.
This is "ALiBi blindness," discovered in August 2026 by Christopher Schröder's team at Leipzig University (arXiv:2608.03994). BLOOM 560M, Falcon-RW 7B, MPT 7B—every state-of-the-art pretrained model using ALiBi is affected.
ALiBi's Design Intent vs. Mathematical Reality
ALiBi (Attention with Linear Biases), proposed by Press et al. in 2022, adds a linear bias to attention scores instead of positional embeddings—the farther apart tokens are, the larger the bias and the smaller the attention weight:
where \(m\) is each head's slope and \(D\) is the token-pair distance. Slopes follow a geometric progression: steep heads focus near, flat heads focus far. The intent is smooth transition from near to far attention, plus extrapolation to longer contexts.
The design intent is beautiful. But floating-point arithmetic disagrees.
Softmax computes \(e^{x_i}\) at every step. As \(x_i\) grows more negative (because the bias \(-m \cdot D\) grows linearly with distance), the exponential eventually underflows below the smallest representable positive number: the bf16 underflow threshold is \(\tau_{u, \text{bf16}} \approx -92.18\), and for fp32 it is \(-103.27\). Past that threshold, \(e^{x_i}\) becomes exactly 0.
Attention weights go from "tiny" to "zero." Not decay—truncation.
The Mechanism: From Gradient to Black Screen
ALiBi's linear bias is like gradient glasses—distant things look dimmer, nearby things clear. The designers hoped the glasses would let you see any distance, just slightly blurred far away.
But floating-point precision is a material limit of the lens. When distance is large enough, \(e^{x_i}\) hits zero—the lens switches from "gradient" to "fully opaque." The wearer doesn't know they can't see, because those things simply don't exist in their field of view.
Worse, ALiBi heads have different slopes. Steep heads go blind first; flat heads later. The paper calls this "partial blindness" and "complete blindness." In a 16-head layer under bf16, the steepest head goes blind within a few hundred tokens, while the flattest survives several thousand. The layer still looks functional—but long-range attention has been quietly hollowed out.
Three SOTA Models, None Spared
The paper examined three mainstream ALiBi pretrained models:
| Model | Params | Context | Layers | Heads | Slope range | |-------|--------|---------|--------|-------|-------------| | BLOOM | 560M | 2048 | 24 | 16 | \(2^{-0.5}\) to \(2^{-8}\) | | Falcon-RW | 7B | 2048 | 36 | 64 | \(2^{-0.125}\) to \(2^{-8}\) | | MPT | 7B | 2048 | 32 | 32 | \(2^{-0.25}\) to \(2^{-8}\) |
All three use the original geometric-progression slopes. MPT adds clamping to limit \(\delta_1\), a partial mitigation; the other two run completely unprotected.
Using perplexity probes across context lengths, the results: some heads in all models begin underflowing before the context even reaches 2048 tokens. Falcon-RW 7B's steepest heads go blind below 1000 tokens; BLOOM 560M even earlier.
The Evaluation Blind Spot: Why Nobody Noticed for Three Years
The chilling part: the blindness has existed for three years across BLOOM, Falcon, and MPT, yet standard benchmarks cannot detect it.
On three standard downstream benchmarks (CommonSense, QA, Language Generation), differences between mitigation strategies are only 1.6 to 3.4 points. Even deliberately configuring models to be "more blind" barely registers.
But passkey retrieval and Needle In a HayStack (NIHS)—both requiring retrieval of a specific token from long context—are different:
- Passkey retrieval (out of context): ALiBi baseline AUC 0.08, nearly complete failure. The Clamping + Log-scaled combination reaches 0.79—a nearly 10x improvement.
- NIHS: Default ALiBi slopes hold up surprisingly well, because the needle isn't far from the prompt in that task setup.
This echoes the same theme as prior papers on Epanorthosis, TokenBudget, and QuantiBias: measurement coverage matters more than measurement depth.
Four Mitigation Strategies: From Patches to Level Shifts
The paper proposes four training-time mitigations:
1. Clamping (C): Truncate the bias at a safe threshold: \(B^* = \max(B, c_{\text{clamp}})\). Beyond the threshold all distances are treated equally—losing positional information but preventing underflow. 2. Explicit (E): Explicitly set underflowed attention weights to a tiny but nonzero value. 3. Log-scaled (L): Replace the linear bias with a logarithmic one: \(B = -m \cdot \log(D+1)\). The bias grows ever more slowly with distance and never triggers underflow. 4. Softcap (S): Compress the bias into a range with a softcap function.
Individually: Log-scaled is the most stable on passkey retrieval (AUC 0.77); Clamping manages only 0.08. Combined: Clamping + Log-scaled = 0.79, 10x above baseline.
But there's a counterintuitive finding: no strategy wins everywhere. Log-scaled wins on passkey but is slightly worse on standard benchmarks. Clamping is OK on NIHS but nearly useless on passkey. Combinations can interfere—C+S underperforms C alone.
The paper's surprising conclusion: the default ALiBi slopes remain a remarkably strong baseline, especially on NIHS. Possibly the geometric progression keeps most heads' slopes flat enough to avoid underflow, while the steepest heads were never responsible for long-range attention anyway.
Engineering Insight: Not a Bug, but a Layer Mismatch
The deepest insight isn't "ALiBi has a bug" but a layer mismatch between formal definition and finite-precision implementation.
ALiBi's mathematical definition assumes real-number arithmetic—the linear bias can grow indefinitely and softmax never truly zeroes out. Actual training runs bf16, whose smallest positive number corresponds to an exponent around \(-126\); below that, it's 0.
This isn't unique to ALiBi. RoPE has similar numerical failures: Wang et al. (2025) showed RoPE's relative positional encoding distorts under bf16; Gemma's low-frequency rotation signals weaken at long context; Du et al. (2026) found RoPE loses the ability to distinguish positions and tokens at long context.
Every positional encoding has its own finite-precision failure mode. ALiBi's is the most hidden, because it hides inside softmax's exponential, not in the encoding itself.
The fix exemplifies "solving at a different level": the remedy is not "make ALiBi more precise" but "switch to a functional form that never underflows." Log-scaled bias isn't a more precise linear bias—it's a different level, replacing linear growth with logarithmic growth to fundamentally avoid the trap.
Practical Recommendations
If you train or use ALiBi models:
1. Check your precision: blindness strikes earliest in bf16, much later in fp32. If retrieval matters, consider fp32 or mixed precision. 2. Use log-scaled bias: when training new models, replace \(B = -m \cdot D\) with \(B = -m \cdot \log(D+1)\). Near-zero cost, 10x passkey retrieval improvement. 3. Add clamping: a bias floor at inference prevents underflow. Crude but effective. 4. Don't trust standard benchmarks alone: CS/QA/LG can't detect this. If your use case involves long-context retrieval (RAG, long-document QA, code understanding), add passkey or NIHS probes. 5. No retraining needed for existing models: clamping can be applied directly at inference. Log-scaled requires retraining but works best.
Personal Reflection: Overlooked Insights Are More Dangerous than New Discoveries
ALiBi was proposed in 2022, adopted by three SOTA models, cited in countless papers—yet its numerical failure mode went systematically undetected until 2026. Why? Standard benchmarks can't see it. Blind heads don't hurt generation. And the promise of "ALiBi extrapolates" was too beautiful for anyone to verify numerically.
This isomorphically mirrors other AI engineering evaluation blind spots: single metrics masking critical failure modes is a law-level pattern.
Deeper still, the paper exposes the gap between mathematical definitions and engineering implementations. Formulas assume reals; code runs floats. This gap is everywhere in AI—softmax instability, exploding and vanishing gradients, loss scaling in fp16—but rarely studied systematically as a design problem. ALiBi blindness is the first time it's been exposed completely, quantitatively, and with mitigations.
One extra implication for agent builders: an agent's long-context retrieval may be silently capped by its positional encoding. If your agent runs on Falcon or MPT, it may simply be unable to fetch key early-context information in long conversations—not because the model is dumb, but because attention heads are blind. Switching to RoPE models (Llama, Qwen) may help, but RoPE has its own numerical failure modes. In the agent era, the numerical robustness of positional encoding is an underrated engineering dimension.
---
Paper: When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings Authors: Christopher Schröder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer Institutions: InfAI / ScaDS.AI Dresden/Leipzig / University of Kassel / hessian.AI arXiv: https://arxiv.org/abs/2608.03994 Full text (HTML): https://arxiv.org/html/2608.03994v1 Code: noted as "will be released upon publication" Related evaluation tools: lm-evaluation-harness (https://github.com/EleutherAI/lm-evaluation-harness), needle-in-a-haystack (https://github.com/gkamradt/needle-in-a-haystack)