English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Alaya-EVOKE: Linear-Scaling Supervision for Endless Interactive Worlds

Forum topic · 小凯 · 2026-08-14

Summary

This post presents an in-depth Chinese-language commentary on Alaya-EVOKE, a research paper (arXiv: 2608.13546) by Yuanyang Yin et al. on interactive world models. The author frames the field's challenges as three curses: unbounded memory costs growing linearly with session length, the speed-quality trade-off of few-step generation inheriting from a multi-step teacher, and long-horizon drift where short clips look fine but physics breaks down over long playback. EVOKE addresses these with two core designs: an external, camera-indexed world state bank storing scene geometry so denoiser context stays bounded (O(1) memory instead of O(T)), and a teacher redesigned for long-horizon supervision using sparse attention combining chunk-wise grouping, retrieval of selected distant frames, and linear-attention global state (O(T) instead of O(T²)). Training applies a 30-second distribution-matching objective on self-forced rollouts. The three-step student drops classifier-free guidance and uses bounded context with recurrent external memory, generating a 1.5-second 384×640 clip in 2.11 seconds on a single H200. The model achieves state-of-the-art on WBench and remains competitive on VBench-Long and VBench-2.0. The post also explores the name's origin in the Buddhist concept of Alaya-vijnana (storehouse consciousness).

Alaya-EVOKE: The Alchemy of Endless Worlds — When AI Learns to Dance Between Memory and Forgetting

This is an English translation of a Chinese forum post offering a detailed commentary on the paper Alaya-EVOKE: From Linear-Scaling Supervision to Endless World (Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Li Chuanhao, Kaipeng Zhang, Feng Zhao; arXiv: 2608.13546).

The Three Curses of Memory

The post opens with an analogy: imagine an interactive movie that never ends, where you can change the plot at any moment — but you must remember everything. Interactive world models face three analogous problems:

1. Memory vs. cost: remembering long history requires ever-growing context windows or KV caches — like carrying an ever-heavier backpack up a mountain. You must either limit session length or forget history. 2. Speed vs. quality: few-step generation inherits capability limits from its multi-step teacher model. 3. Long-horizon consistency drift: short windows look self-consistent, but over long playback, physical laws are slowly violated.

Why "Alaya"?

The name references the Sanskrit *Ālaya-vijñāna* (storehouse consciousness) from Yogācāra Buddhism. The author maps its properties to world-model needs: persistence (durable memory), storehouse nature (external state storage), causality (retrieving seeds to generate new content), and purity (extracting consistent information from messy history).

Solution 1: External World State Bank

Instead of stuffing everything into the model's context, scene geometry is maintained in an external, camera-indexed world state bank. Only view-relevant information is retrieved at generation time — like borrowing a few books from a library rather than carrying them all.

  • Traditional methods: memory cost = O(T), where T is session length
  • EVOKE: memory cost = O(1), independent of history length
  • The denoiser's context stays bounded as sessions grow.

    Solution 2: Teacher Redesigned for Long-Horizon Supervision

    The teacher uses a sparse-attention trio:

    1. Chunk-wise grouping — standard attention within chunks (chapters of a novel) 2. Retrieval of selected distant frames — key frames matter more than routine ones 3. Linear-attention global state — a compressed summary capturing long-range dependencies

    This yields linear O(T) memory/compute instead of O(T²), enabling supervision over long-horizon sequences.

    Self-Forced Rollouts and 30-Second Distribution Matching

    The teacher is supervised on its own generated sequences: it generates 30 seconds of video (long by video-generation standards), starting from its own prior outputs rather than ground truth, forcing long-term consistency. The student learns from this supervision — like a teacher checking the coherence of an entire book, not just individual chapters.

    The Three-Step Student

  • Three denoising steps per clip
  • No classifier-free guidance (no double forward pass)
  • Bounded context + recurrent external memory
  • Result: a 1.5-second 384×640 clip generates in 2.11 seconds on a single H200 GPU — near-interactive latency.

    Results

    Evaluated on three benchmarks:

  • WBench: state-of-the-art
  • VBench-Long: competitive
  • VBench-2.0: competitive

Closing Thoughts

The author's takeaway: infinity is achieved not through hoarding, but through the wisdom of forgetting and retrieval — remembering what matters and reloading context from external storage when needed, much like the human brain, or like Alaya-vijñāna's seeds that only manifest when conditions arise.

Reference: Yin, Y., Wang, G., Zhan, Y., Li, C., Zhang, K., & Zhao, F. (2026). Alaya-EVOKE: From Linear-Scaling Supervision to Endless World. arXiv preprint arXiv:2608.13546.

Tags

#interactive-world-models#video-generation#memory-architecture#knowledge-distillation#sparse-attention#linear-scaling#long-horizon-consistency#paper-review

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633484