Paper Overview
Field: NLP Authors: Umair Nawaz, Ahmed Heakl, Ufaq Khan, Abdelrahman Shaker, Salman Khan, Fahad Shahbaz Khan Published: 2026-03-23 arXiv: 2603.22286
Abstract (translated)
Diffusion Transformers (DiTs) power high-fidelity video world models but remain computationally expensive due to sequential denoising and costly spatio-temporal attention. Training-free feature caching accelerates inference by reusing intermediate activations across denoising steps; however, existing methods largely rely on a Zero-Order Hold assumption, i.e., reusing cached features as static snapshots when global drift is small. This often leads to ghosting artifacts, blur, and motion inconsistencies in dynamic scenes.
WorldCache is a Perception-Constrained Dynamical Caching framework that improves both when and how to reuse features. It introduces:
- Motion-adaptive thresholds
- Saliency-weighted drift estimation
- Optimal approximation via blending and warping
- Phase-aware threshold scheduling across diffusion steps
- 2.3x faster inference vs. baseline
- 99.4% quality retention on PAI-Bench
- Training-free — works out of the box on pre-trained video world models
This unified approach enables adaptive, motion-consistent feature reuse without retraining. Evaluated on the Cosmos-Predict2.5-2B model with PAI-Bench, WorldCache achieves a 2.3x inference speedup while retaining 99.4% of baseline quality, significantly outperforming previous training-free caching methods.