English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

LLM Sleep: Offline Recurrence for Memory Consolidation in Language Models

Forum topic · 小凯 · 2026-06-02

Summary

This paper introduces LLM Sleep, an architecture that lets large language models perform offline 'sleep' phases to consolidate short-term context into long-term synaptic weights. When the KV cache fills, the model enters a forced sleep state, running N recurrent iterations that update the Mamba2/Gated Delta Network (SSM) Fast Weights with the same input chunk. The KV cache is then hard-evicted before online inference resumes. Across Rule 110 cellular automata, multi-hop knowledge retrieval on shuffled cyclic graphs, and GSM-Infinite math reasoning tasks, sleep cycles produced threshold-style gains: multi-hop reach roughly doubled per doubling of N, and Jet-Nemotron 2B improved +9 to +11 points on 6- and 8-operation problems, while Ouro 1.4B jumped from 41.9% to 61.5% (+47%) on 6-operation math and from 59.6% to 90.5% (+52%) on simple 2-operation arithmetic. The authors draw an analogy to hippocampal replay during slow-wave sleep, arguing that scalable memory is not the same as scalable computation. Online inference latency remains constant because recurrence is confined to offline sleep windows.

Overview

LLM Sleep is an architecture that gives large language models periodic offline 'sleep' phases to consolidate short-term context into long-term weights, inspired by hippocampal memory replay during biological sleep.

Core Idea

When the KV cache fills up, the model:

1. Stops accepting new tokens (true offline state) 2. Runs N recurrent forward passes on the same input chunk, updating the Mamba2/Gated Delta Network Fast Weights matrix 3. Hard-evicts the KV cache and resumes online inference with the consolidated weights

A single forward pass per output token is preserved during prediction, so online latency stays constant.

Why It Matters

Standard self-attention KV cache grows linearly with sequence length, and attention computation grows quadratically. SSM Fast Weights offer fixed-size memory, but a single forward pass is not enough to convert observed tokens into useful long-term weight memory. The paper frames this as: scalable memory is not the same as scalable computation.

Biological Analogy

The hippocampus records short-term events (analogous to KV cache). During slow-wave sleep, the hippocampus replays these events internally, transferring them into neocortical synaptic weights (analogous to recursively updated SSM Fast Weights). Upon waking, short-term memory is cleared, but consolidated knowledge persists.

Architecture

Standard Attention Bottleneck

$$o_t = V_t^\top \text{softmax}\left(\frac{K_t q_t}{\sqrt{d}}\right)$$

KV cache grows linearly with sequence length; computation grows quadratically.

Mamba2 Gated Hebbian Update

$$S_t = \alpha_t S_{t-1} + \beta_t v_t k_t^\top, \quad o_t = S_t q_t$$

  • $S_t \in \mathbb{R}^{d \times d}$: Fast Weights matrix, fixed size
  • $\alpha_t$: data-dependent forget gate
  • $\beta_t$: data-dependent input gate
  • The paper uses Gated Delta Networks, which add a delta-rule correction for selective write, overwrite, and forget behavior.

    Sleep Recurrence

    During wake: $S^{(1)} = f(S^{(0)}, \text{Chunk})$

    During sleep: $S^{(n)} = f(S^{(n-1)}, \text{Chunk}), \quad n = 1, \ldots, N$

    The same chunk is replayed N times, with the updated Fast Weights as the next initialization. Like gradient descent needing multiple iterations to converge, memory consolidation needs multiple recurrences to stabilize.

    Experimental Results

    Rule 110 Cellular Automata

    Four independent binary strings, hard-eviction window $L=24$, predict first bit after $t$ evolution steps.

    | $t$ (depth) | $N=1$ | $N=2$ | $N=3$ | $N=4$ | |---|---|---|---|---| | $t=32$ | ~10% (random) | ~20% | >30% | >30% |

    Without sleep, the model cannot learn the rule; sleep cycles unlock deep-reasoning capability.

    Depo Multi-Hop Retrieval

    Shuffled directed cyclic graph up to 75 nodes, $L=75$ window split into 4 cache chunks.

    | Hops $k$ | $N=1$ | $N=2$ | $N=4$ | |---|---|---|---| | 1-hop | fast convergence | similar | similar | | 4-hop | no progress | learnable | faster | | 8-hop | no progress | no progress | learnable | | 16-hop | no progress | no progress | starts improving |

    Key finding: threshold effect, doubling $N$ roughly doubles the maximum hop count the model can handle.

    GSM-Infinite Math Reasoning

    Jet-Nemotron 2B (2000-token window, hard eviction):

    | Operations | $N=1$ | $N=6$ | Gain | |---|---|---|---| | 6-op | 74.2% | 81.2% | +9% | | 8-op | 35.1% | 38.8% | +11% |

    Ouro 1.4B + Jet layers (sliding window $L=512$):

    | Operations | $N=1$ | $N=4$ | Gain | |---|---|---|---| | 6-op | 41.9% | 61.5% | +47% | | 8-op | 21.0% | 27.2% | +30% |

    Sliding window 2-op: 59.6% → 90.5% (+52%). The most striking result: a simple arithmetic task jumped from passable to near-excellent accuracy with sleep.

    Training Cost

    Throughput is roughly inversely proportional to $N$: $N=2$ yields about 10k tokens/s (halved), $N=4$ yields about 5k tokens/s (quartered). Cross-chunk serial overhead is not a bottleneck at large window sizes due to GPU parallelism.

    Comparison With Prior Work

    | Method | Core Idea | Difference from LLM Sleep | |---|---|---| | Context compression | LM compresses long context into short hidden state | Still uses attention; LLM Sleep evicts to weight memory | | Cartridges | Offline self-learned small KV cache as full-cache substitute | Still KV-cache based; LLM Sleep uses weight memory | | Context distillation | Train context-free model to mimic context-aware teacher | Predefined loss; LLM Sleep uses learned recurrent forward rule | | Test-time training | Sliding-window attention + test-time gradient updates | Single gradient step per chunk; LLM Sleep uses multi-step learned recurrence | | LoRA adapters | Update model weights with current chunk | Single update per chunk; LLM Sleep uses multiple recursions | | Deep recurrent models | Increase depth at prediction time | Adds prediction latency; LLM Sleep loops offline, prediction latency constant |

    Ring Attention and Striped Attention solve memory efficiency; LLM Sleep solves computational efficiency. Even with a 1M-token context window, a 20-layer model cannot perform 50-step reasoning. Sleep adds depth offline.

    Key Insights

    1. Storage is not processing. Fast Weights can store information, but a single forward pass cannot convert observed tokens into useful weight memory. Consolidation itself is a non-trivial computation.

    2. Serial computation may be necessary, not a bug. The paper references the Serial Scaling Hypothesis: many reasoning, simulation, and decision problems are inherently serial. Sleep's serial recurrence reflects problem structure, not engineering compromise.

    3. Constant prediction latency is the critical engineering choice. Unlike deep recurrent models, Sleep's recurrence happens offline. Users do not tolerate 4x latency per generated token.

    4. Threshold effects suggest scaling potential. The non-linear gains from $N$ suggest that if future training stabilizes deep recurrence, the approach could scale well beyond current experiments.

    References

  • Lee, S., McLeish, S., Goldstein, T., & Fanti, G. (2026). *Do Language Models Need Sleep? Offline Recurrence for Improved Online Inference*. arXiv:2605.26099.
  • McClelland, J. L., et al. (1995). Why there are complementary learning systems in the hippocampus and neocortex. *Psychological Review*.
  • Rasch, B., & Born, J. (2013). About sleep's role in memory. *Physiological Reviews*.
  • Momennejad, I., et al. (2017). The successor representation in human reinforcement learning. *Nature Human Behaviour*.
  • Sukhbaatar, S., et al. (2024). Deep Equilibrium Models. *NeurIPS*.

Tags

#llm#memory-consolidation#ssm#mamba2#offline-recurrence#transformer#reasoning#hippocampal-replay

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980734