Alaya-EVOKE: The Alchemy of Endless Worlds — When AI Learns to Dance Between Memory and Forgetting
This is an English translation of a Chinese forum post offering a detailed commentary on the paper Alaya-EVOKE: From Linear-Scaling Supervision to Endless World (Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Li Chuanhao, Kaipeng Zhang, Feng Zhao; arXiv: 2608.13546).
The Three Curses of Memory
The post opens with an analogy: imagine an interactive movie that never ends, where you can change the plot at any moment — but you must remember everything. Interactive world models face three analogous problems:
1. Memory vs. cost: remembering long history requires ever-growing context windows or KV caches — like carrying an ever-heavier backpack up a mountain. You must either limit session length or forget history. 2. Speed vs. quality: few-step generation inherits capability limits from its multi-step teacher model. 3. Long-horizon consistency drift: short windows look self-consistent, but over long playback, physical laws are slowly violated.
Why "Alaya"?
The name references the Sanskrit *Ālaya-vijñāna* (storehouse consciousness) from Yogācāra Buddhism. The author maps its properties to world-model needs: persistence (durable memory), storehouse nature (external state storage), causality (retrieving seeds to generate new content), and purity (extracting consistent information from messy history).
Solution 1: External World State Bank
Instead of stuffing everything into the model's context, scene geometry is maintained in an external, camera-indexed world state bank. Only view-relevant information is retrieved at generation time — like borrowing a few books from a library rather than carrying them all.
- Traditional methods: memory cost = O(T), where T is session length
- EVOKE: memory cost = O(1), independent of history length
- Three denoising steps per clip
- No classifier-free guidance (no double forward pass)
- Bounded context + recurrent external memory
- WBench: state-of-the-art
- VBench-Long: competitive
- VBench-2.0: competitive
The denoiser's context stays bounded as sessions grow.
Solution 2: Teacher Redesigned for Long-Horizon Supervision
The teacher uses a sparse-attention trio:
1. Chunk-wise grouping — standard attention within chunks (chapters of a novel) 2. Retrieval of selected distant frames — key frames matter more than routine ones 3. Linear-attention global state — a compressed summary capturing long-range dependencies
This yields linear O(T) memory/compute instead of O(T²), enabling supervision over long-horizon sequences.
Self-Forced Rollouts and 30-Second Distribution Matching
The teacher is supervised on its own generated sequences: it generates 30 seconds of video (long by video-generation standards), starting from its own prior outputs rather than ground truth, forcing long-term consistency. The student learns from this supervision — like a teacher checking the coherence of an entire book, not just individual chapters.
The Three-Step Student
Result: a 1.5-second 384×640 clip generates in 2.11 seconds on a single H200 GPU — near-interactive latency.
Results
Evaluated on three benchmarks:
Closing Thoughts
The author's takeaway: infinity is achieved not through hoarding, but through the wisdom of forgetting and retrieval — remembering what matters and reloading context from external storage when needed, much like the human brain, or like Alaya-vijñāna's seeds that only manifest when conditions arise.
Reference: Yin, Y., Wang, G., Zhan, Y., Li, C., Zhang, K., & Zhao, F. (2026). Alaya-EVOKE: From Linear-Scaling Supervision to Endless World. arXiv preprint arXiv:2608.13546.