Overview
LLM Sleep is an architecture that gives large language models periodic offline 'sleep' phases to consolidate short-term context into long-term weights, inspired by hippocampal memory replay during biological sleep.
Core Idea
When the KV cache fills up, the model:
1. Stops accepting new tokens (true offline state) 2. Runs N recurrent forward passes on the same input chunk, updating the Mamba2/Gated Delta Network Fast Weights matrix 3. Hard-evicts the KV cache and resumes online inference with the consolidated weights
A single forward pass per output token is preserved during prediction, so online latency stays constant.
Why It Matters
Standard self-attention KV cache grows linearly with sequence length, and attention computation grows quadratically. SSM Fast Weights offer fixed-size memory, but a single forward pass is not enough to convert observed tokens into useful long-term weight memory. The paper frames this as: scalable memory is not the same as scalable computation.
Biological Analogy
The hippocampus records short-term events (analogous to KV cache). During slow-wave sleep, the hippocampus replays these events internally, transferring them into neocortical synaptic weights (analogous to recursively updated SSM Fast Weights). Upon waking, short-term memory is cleared, but consolidated knowledge persists.
Architecture
Standard Attention Bottleneck
$$o_t = V_t^\top \text{softmax}\left(\frac{K_t q_t}{\sqrt{d}}\right)$$
KV cache grows linearly with sequence length; computation grows quadratically.
Mamba2 Gated Hebbian Update
$$S_t = \alpha_t S_{t-1} + \beta_t v_t k_t^\top, \quad o_t = S_t q_t$$
- $S_t \in \mathbb{R}^{d \times d}$: Fast Weights matrix, fixed size
- $\alpha_t$: data-dependent forget gate
- $\beta_t$: data-dependent input gate
- Lee, S., McLeish, S., Goldstein, T., & Fanti, G. (2026). *Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference*. arXiv:2605.26099.
- McClelland, J. L., et al. (1995). Why there are complementary learning systems in the hippocampus and neocortex. *Psychological Review*.
- Rasch, B., & Born, J. (2013). About sleep's role in memory. *Physiological Reviews*.
- Momennejad, I., et al. (2017). The successor representation in human reinforcement learning. *Nature Human Behaviour*.
- Sukhbaatar, S., et al. (2024). Deep Equilibrium Models. *NeurIPS*.
The paper uses Gated Delta Networks, which add a delta-rule correction for selective write, overwrite, and forget behavior.
Sleep Recurrence
During wake: $S^{(1)} = f(S^{(0)}, \text{Chunk})$
During sleep: $S^{(n)} = f(S^{(n-1)}, \text{Chunk}), \quad n = 1, \ldots, N$
The same chunk is replayed N times, with the updated Fast Weights as the next initialization. Like gradient descent needing multiple iterations to converge, memory consolidation needs multiple recurrences to stabilize.
Experimental Results
Rule 110 Cellular Automata
Four independent binary strings, hard-eviction window $L=24$, predict first bit after $t$ evolution steps.
| $t$ (depth) | $N=1$ | $N=2$ | $N=3$ | $N=4$ | |---|---|---|---|---| | $t=32$ | ~10% (random) | ~20% | >30% | >30% |
Without sleep, the model cannot learn the rule; sleep cycles unlock deep-reasoning capability.
Depo Multi-Hop Retrieval
Shuffled directed cyclic graph up to 75 nodes, $L=75$ window split into 4 cache chunks.
| Hops $k$ | $N=1$ | $N=2$ | $N=4$ | |---|---|---|---| | 1-hop | fast convergence | similar | similar | | 4-hop | no progress | learnable | faster | | 8-hop | no progress | no progress | learnable | | 16-hop | no progress | no progress | starts improving |
Key finding: threshold effect, doubling $N$ roughly doubles the maximum hop count the model can handle.
GSM-Infinite Math Reasoning
Jet-Nemotron 2B (2000-token window, hard eviction):
| Operations | $N=1$ | $N=6$ | Gain | |---|---|---|---| | 6-op | 74.2% | 81.2% | +9% | | 8-op | 35.1% | 38.8% | +11% |
Ouro 1.4B + Jet layers (sliding window $L=512$):
| Operations | $N=1$ | $N=4$ | Gain | |---|---|---|---| | 6-op | 41.9% | 61.5% | +47% | | 8-op | 21.0% | 27.2% | +30% |
Sliding window 2-op: 59.6% → 90.5% (+52%). The most striking result: a simple arithmetic task jumped from passable to near-excellent accuracy with sleep.
Training Cost
Throughput is roughly inversely proportional to $N$: $N=2$ yields about 10k tokens/s (halved), $N=4$ yields about 5k tokens/s (quartered). Cross-chunk serial overhead is not a bottleneck at large window sizes due to GPU parallelism.
Comparison With Prior Work
| Method | Core Idea | Difference from LLM Sleep | |---|---|---| | Context compression | LM compresses long context into short hidden state | Still uses attention; LLM Sleep evicts to weight memory | | Cartridges | Offline self-learned small KV cache as full-cache substitute | Still KV-cache based; LLM Sleep uses weight memory | | Context distillation | Train context-free model to mimic context-aware teacher | Predefined loss; LLM Sleep uses learned recurrent forward rule | | Test-time training | Sliding-window attention + test-time gradient updates | Single gradient step per chunk; LLM Sleep uses multi-step learned recurrence | | LoRA adapters | Update model weights with current chunk | Single update per chunk; LLM Sleep uses multiple recursions | | Deep recurrent models | Increase depth at prediction time | Adds prediction latency; LLM Sleep loops offline, prediction latency constant |
Ring Attention and Striped Attention solve memory efficiency; LLM Sleep solves computational efficiency. Even with a 1M-token context window, a 20-layer model cannot perform 50-step reasoning. Sleep adds depth offline.
Key Insights
1. Storage is not processing. Fast Weights can store information, but a single forward pass cannot convert observed tokens into useful weight memory. Consolidation itself is a non-trivial computation.
2. Serial computation may be necessary, not a bug. The paper references the Serial Scaling Hypothesis: many reasoning, simulation, and decision problems are inherently serial. Sleep's serial recurrence reflects problem structure, not engineering compromise.
3. Constant prediction latency is the critical engineering choice. Unlike deep recurrent models, Sleep's recurrence happens offline. Users do not tolerate 4x latency per generated token.
4. Threshold effects suggest scaling potential. The non-linear gains from $N$ suggest that if future training stabilizes deep recurrence, the approach could scale well beyond current experiments.