English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Alaya-EVOKE: A World Model That Generates Endless, Interactive Worlds With Linear-Scaling Supervision

Forum topic · 小凯 · 2026-08-15

Summary

Alaya-EVOKE is an interactive video world model designed for persistent memory, responsive interaction, and open-ended long-horizon generation. It addresses two core limitations of existing models: the growing cost of maintaining history in denoiser contexts, and the capability ceiling of few-step generation. EVOKE externalizes persistent world state into a camera-indexed world state bank, retrieving only view-relevant information so the denoiser context stays bounded as sessions grow. Its teacher model uses sparse attention combining chunk-wise grouping, retrieval of distant frames, and a linear-attention global state, achieving linear memory and compute growth for long-horizon supervision. A 30-second distribution-matching objective under self-forced rollouts distills these capabilities into a three-step student model that requires no classifier-free guidance. On a single H200 GPU at 384x640 resolution, each 1.5-second chunk generates in 2.11 seconds. EVOKE achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.

Alaya-EVOKE: A World Model That Generates Endless, Interactive Worlds

> *"Reality is merely a shared hallucination—and now we're teaching machines to weave their own: infinite in length, detail, and possibility."*

Alaya-EVOKE (EVOKE) is an interactive world model that can generate high-quality video segments, seamlessly chained into a "never-ending" world. The original forum post is a long-form Chinese commentary on the paper; below is a full English rendering of its core technical content.

Paper Information

  • Title: Alaya-EVOKE: From Linear-Scaling Supervision to Endless World
  • Authors: Yuanyang Yin, Gongxuan Wang, Yifan Zhan, Chuanhao Li, Kaipeng Zhang, Feng Zhao
  • Category: cs.CV
  • arXiv ID: 2608.13546v1
  • Published: 2026-08-15
  • Abstract

    Interactive world models must support persistent memory, responsive interaction, and long-horizon generation, yet these requirements place conflicting demands on the model. Maintaining history in the denoiser context or key-value cache incurs growing cost, forcing a trade-off between session length and retained memory, while low-latency interaction relies on few-step generation whose capabilities are bounded by its teacher. Evoke addresses both limitations by externalizing persistent world state and redesigning the teacher for long-horizon interactive generation. Scene geometry is maintained in an external, camera-indexed world state bank, from which only view-relevant information is retrieved, keeping the denoiser context bounded as session grows. Rather than treating the teacher as a fixed generator, we design it for long-horizon supervision: its sparse attention combines chunk-wise grouping, retrieval of selected distant frames, and a linear-attention global state, yielding linear growth in memory and compute while enabling supervision over long horizons. Such supervision exposes content drift that stays locally plausible within short windows, while per-chunk conditioning enables prompt changes and event control throughout the sequence. A 30-second distribution-matching objective, applied under self-forced rollouts, transfers both capabilities to a three-step student that uses no classifier-free guidance, improving resistance to long-term drift while preserving responsive conditioning. With bounded context and recurrent external memory, Evoke supports open-ended, continuously evolving generation; on a single H200 at \(384\times 640\), each \(1.5\,\mathrm{s}\) chunk is generated in \(2.11\,\mathrm{s}\). As a three-step world model, Evoke achieves state-of-the-art performance on WBench while remaining competitive on VBench-Long and VBench-2.0.

    The Three Paradoxes of World Models

    1. Memory vs. Length

    Most world models use Transformer self-attention: each new frame must attend to all previous outputs. This makes memory and compute grow quadratically with video length — like a reader who must keep every page of a 1000-page novel spread open on the desk. EVOKE's first goal is to teach this reader to "take notes" — extract key information rather than hold everything in context.

    2. Interaction vs. Quality

    Generation quality depends on many denoising steps, but responsive interaction demands few-step generation. It's the painter's dilemma: paint fast and crude, or paint beautifully while the subject has long since changed pose. EVOKE aims for a painter who is both precise and fast.

    3. Short-Term Coherence vs. Long-Term Drift

    Even with long videos and fast interaction, models suffer from content drift: after a minute of generation, textures change mid-scene, clocks disagree, and faces subtly morph into someone else — because each frame only sees a limited recent window. EVOKE seeks to make the machine remember the world's "essence," not just its latest impression.

    EVOKE's Two Keys

    Key 1: The External Memory Palace — World State Bank

    EVOKE stores scene geometry (3D structure, depth, surface appearance) in an external, camera-indexed world state bank — like a director's script book for a 10-hour epic. When generating frame 100, the model doesn't "recall" early frames; it queries: *which region of the 3D world does the current camera view correspond to, and what is its structure?*

    Crucially, this memory is bounded: the world's map is finite (a museum is only so big), so the model stores the world's structure rather than every frame's pixels. The denoiser context stays bounded no matter how long the session runs.

    Key 2: The Teacher with "Clairvoyance" — Long-Horizon Supervision

    Existing models train like teaching a child to bike by only watching the next few meters — they never learn to "ride the whole street." EVOKE's teacher is designed for long-horizon supervision with three components:

    1. Chunk-wise grouping: fine-grained processing within local chunks 2. Distant frame retrieval: actively "recalling" key far-away frames 3. Linear-attention global state: a compressed summary of the video's overall trajectory

    Together, these give the teacher linear O(n) memory and compute instead of O(n²), enabling supervision over arbitrarily long videos and exposing content drift that would look locally plausible in short windows.

    The Three-Step Student

    The teacher is too slow for interaction, so EVOKE distills it into a student that generates in 3 steps with no classifier-free guidance. Training uses a 30-second distribution-matching objective under self-forced rollouts: the model generates 30 seconds of video and is checked for consistency — do floor textures drift? Does the character's appearance change? The student learns both fast generation and long-term consistency.

    Capabilities

    Unbounded coherent generation — With external recurrent memory, EVOKE generates continuously without hitting a "memory wall." On a single H200, each 1.5s chunk takes 2.11s — nearly real-time.

    Responsive interaction — The 3-step student reacts instantly to instructions ("turn left," "make it rain") while staying consistent with prior content.

    Event-driven narrative — Per-chunk conditioning lets users assign different prompts to different segments ("sunny morning" → "sudden downpour" → "rainbow after rain"), with natural transitions like darkening skies and gradually drying clothes.

    Benchmark Results

  • WBench: State-of-the-art, notably on geometry consistency and interaction fidelity
  • VBench-Long: Competitive, with strong temporal coherence
  • VBench-2.0: Competitive without sacrificing short-horizon visual quality
| Metric | Value | |---|---| | Generation speed | 2.11s per 1.5s chunk (H200, 384×640) | | Teacher attention complexity | O(n) linear vs. O(n²) | | Student sampling steps | 3 (vs. 50+ typical) |

Applications and Reflections

The post highlights potential applications: infinitely expanding open-world games, interactive films that remember viewer choices, virtual city simulation for urban planning, and scientific visualization (walking through a cell, flying through a galaxy).

It also raises risks: harder-to-detect deepfakes in a world of coherent synthetic realities, over-immersion in compelling virtual worlds, and ambiguous authorship of AI-generated worlds and their cultures.

Reference

Yin, Y., Wang, G., Zhan, Y., Li, C., Zhang, K., & Zhao, F. (2026). Alaya-EVOKE: From Linear-Scaling Supervision to Endless World. *arXiv preprint arXiv:2608.13546*.

Tags

#world-models#video-generation#alaya-evoke#long-horizon-consistency#diffusion-models#arxiv#ai-research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178633540