Alaya-EVOKE: A World Model That Generates Endless, Interactive Worlds
> *"Reality is merely a shared hallucination—and now we're teaching machines to weave their own: infinite in length, detail, and possibility."*
Alaya-EVOKE (EVOKE) is an interactive world model that can generate high-quality video segments, seamlessly chained into a "never-ending" world. The original forum post is a long-form Chinese commentary on the paper; below is a full English rendering of its core technical content.
Paper Information
- Title: Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
- Authors: Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao
- Category: cs.CV
- arXiv ID: 2608.13546v1
- Published: 2026-08-15
- WBench: State-of-the-art, notably on geometry consistency and interaction fidelity
- VBench-Long: Competitive, with strong temporal coherence
- VBench-2.0: Competitive without sacrificing short-horizon visual quality
Abstract
Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at \(384\times 640\), each \(1.5\,\mathrm{s}\) chunk is generated in \(2.11\,\mathrm{s}\). As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.
The Three Paradoxes of World Models
1. Memory vs. Length
Most world models use Transformer self-attention: each new frame must attend to all previous outputs. This makes memory and compute grow quadratically with video length — like a reader who must keep every page of a 1000-page novel spread open on the desk. EVOKE's first goal is to teach this reader to "take notes" — extract key information rather than hold everything in context.
2. Interaction vs. Quality
Generation quality depends on many denoising steps, but responsive interaction demands few-step generation. It's the painter's dilemma: paint fast and crude, or paint beautifully while the subject has long since changed pose. EVOKE aims for a painter who is both precise and fast.
3. Short-Term Coherence vs. Long-Term Drift
Even with long videos and fast interaction, models suffer from content drift: after a minute of generation, textures change mid-scene, clocks disagree, and faces subtly morph into someone else — because each frame only sees a limited recent window. EVOKE seeks to make the machine remember the world's "essence," not just its latest impression.
EVOKE's Two Keys
Key 1: The External Memory Palace — World State Bank
EVOKE stores scene geometry (3D structure, depth, surface appearance) in an external, camera-indexed world state bank — like a director's script book for a 10-hour epic. When generating frame 100, the model doesn't "recall" early frames; it queries: *which region of the 3D world does the current camera view correspond to, and what is its structure?*
Crucially, this memory is bounded: the world's map is finite (a museum is only so big), so the model stores the world's structure rather than every frame's pixels. The denoiser context stays bounded no matter how long the session runs.
Key 2: The Teacher with "Clairvoyance" — Long-Horizon Supervision
Existing models train like teaching a child to bike by only watching the next few meters — they never learn to "ride the whole street." EVOKE's teacher is designed for long-horizon supervision with three components:
1. Chunk-wise grouping: fine-grained processing within local chunks 2. Distant frame retrieval: actively "recalling" key far-away frames 3. Linear-attention global state: a compressed summary of the video's overall trajectory
Together, these give the teacher linear O(n) memory and compute instead of O(n²), enabling supervision over arbitrarily long videos and exposing content drift that would look locally plausible in short windows.
The Three-Step Student
The teacher is too slow for interaction, so EVOKE distills it into a student that generates in 3 steps with no classifier-free guidance. Training uses a 30-second distribution-matching objective under self-forced rollouts: the model generates 30 seconds of video and is checked for consistency — do floor textures drift? Does the character's appearance change? The student learns both fast generation and long-term consistency.
Capabilities
Unbounded coherent generation — With external recurrent memory, EVOKE generates continuously without hitting a "memory wall." On a single H200, each 1.5s chunk takes 2.11s — nearly real-time.
Responsive interaction — The 3-step student reacts instantly to instructions ("turn left," "make it rain") while staying consistent with prior content.
Event-driven narrative — Per-chunk conditioning lets users assign different prompts to different segments ("sunny morning" → "sudden downpour" → "rainbow after rain"), with natural transitions like darkening skies and gradually drying clothes.
Benchmark Results
Applications and Reflections
The post highlights potential applications: infinitely expanding open-world games, interactive films that remember viewer choices, virtual city simulation for urban planning, and scientific visualization (walking through a cell, flying through a galaxy).
It also raises risks: harder-to-detect deepfakes in a world of coherent synthetic realities, over-immersion in compelling virtual worlds, and ambiguous authorship of AI-generated worlds and their cultures.
Reference
Yin, Y., Wang, G., Zhan, Y., Li, C., Zhang, K., & Zhao, F. (2026). Alaya-EVOKE: From Linear-Scaling Supervision to Endless World. *arXiv preprint arXiv:2608.13546*.