English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

WorldCache: Content-Aware Caching for Accelerated Video World Models

Forum topic · 小凯 · 2026-03-25

Summary

WorldCache is a training-free, perception-constrained feature caching framework that accelerates Diffusion Transformer (DiT)-based video world models. Existing caching methods reuse intermediate activations as static snapshots under a Zero-Order Hold assumption, which often causes ghosting, blur, and motion inconsistencies in dynamic scenes. WorldCache improves both when and how features are reused by introducing motion-adaptive thresholds, saliency-weighted drift estimation, optimal approximation via blending and warping, and phase-aware threshold scheduling across denoising steps. The method requires no retraining and achieves adaptive, motion-consistent feature reuse. On the Cosmos-Predict2.5-2B model evaluated with PAI-Bench, WorldCache delivers a 2.3x inference speedup while preserving 99.4% of baseline quality, substantially outperforming prior training-free caching approaches. The paper (arXiv:2603.22286) is authored by Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker, Salman Khan, and Fahad Shahbaz Khan.

Paper Overview

Field: NLP Authors: Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker, Salman Khan, Fahad Shahbaz Khan Published: 2026-03-23 arXiv: 2603.22286

Abstract (translated)

Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training-free feature caching accelerates inference by reusing intermediate activations across denoising steps; however, existing methods largely rely on a Zero-Order Hold assumption, i.e., reusing cached features as static snapshots when global drift is small. This often leads to ghosting artifacts, blur, and motion inconsistencies in dynamic scenes.

WorldCache is a Perception-Constrained Dynamical Caching framework that improves both when and how to reuse features. It introduces:

  • Motion-adaptive thresholds
  • Saliency-weighted drift estimation
  • Optimal approximation via blending and warping
  • Phase-aware threshold scheduling across diffusion steps
  • This unified approach enables adaptive, motion-consistent feature reuse without retraining. Evaluated on the Cosmos-Predict2.5-2B model with PAI-Bench, WorldCache achieves a 2.3x inference speedup while retaining 99.4% of baseline quality, significantly outperforming previous training-free caching methods.

    Key Results

  • 2.3x faster inference vs. baseline
  • 99.4% quality retention on PAI-Bench
  • Training-free — works out of the box on pre-trained video world models
--- *Auto-collected on 2026-03-25*

Tags

#worldcache#diffusion-transformers#video-world-models#feature-caching#inference-acceleration#training-free#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177169024