English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

YoCausal: Measuring How Far Video Generation Models Are from True World Models via Causality Evaluation

Forum topic · 小凯 · 2026-06-01

Summary

YoCausal is a two-level benchmark designed to test whether video diffusion models (VDMs) truly understand causality or merely overfit to statistical temporal patterns. Inspired by the violation-of-expectation (VoE) paradigm from cognitive science, the benchmark creates natural counterfactual samples at zero cost by time-reversing real-world videos, yielding a freely scalable evaluation protocol. Level 1 introduces the Reverse Surprise Index (RSI), which quantifies a model's perception of temporal direction via denoising loss. Level 2 introduces the Causal Cognition Index (CCI), which uses vision-language models to partition datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias. Evaluations of 13 state-of-the-art VDMs reveal that perceiving the arrow of time does not imply an understanding of causality, and that a significant gap remains relative to human-level causal cognition. The work was posted on zhichai.net with arXiv identifier 2605.30346 by researchers including You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, and Zhixiang Wang.

Paper Overview

Field: Computer Vision (CV) Authors: You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, Zhixiang Wang Posted: 2026-05-28 arXiv: 2605.30346

Summary

As video diffusion models (VDMs) advance toward becoming world models, a key question arises: do they truly understand causality, or are they merely overfitting to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data, and the sim-to-real gap limits real-world generalization.

The authors propose YoCausal, a two-level benchmark inspired by the violation-of-expectation (VoE) paradigm from cognitive science. By time-reversing real-world videos at zero cost to serve as natural counterfactual samples, YoCausal establishes a freely scalable evaluation protocol.

  • Level 1 — Reverse Surprise Index (RSI): quantifies a model's perception of the temporal arrow via denoising loss.
  • Level 2 — Causal Cognition Index (CCI): uses vision-language models to stratify datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias.

Findings

Evaluations across 13 state-of-the-art VDMs reveal that perceiving the arrow of time does not imply an understanding of causality, and a significant gap remains relative to human-level causal cognition.

---

*Auto-collected on 2026-06-01*

Tags

#video-diffusion-models#world-models#causality#benchmark#computer-vision#evaluation#arxiv#cognitive-science

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177980673