Summary
YoCausal is a two-level benchmark designed to test whether video diffusion models (VDMs) truly understand causality or merely overfit to statistical temporal patterns. Inspired by the violation-of-expectation (VoE) paradigm from cognitive science, the benchmark creates natural counterfactual samples at zero cost by time-reversing real-world videos, yielding a freely scalable evaluation protocol. Level 1 introduces the Reverse Surprise Index (RSI), which quantifies a model's perception of temporal direction via denoising loss. Level 2 introduces the Causal Cognition Index (CCI), which uses vision-language models to partition datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias. Evaluations of 13 state-of-the-art VDMs reveal that perceiving the arrow of time does not imply an understanding of causality, and that a significant gap remains relative to human-level causal cognition. The work was posted on zhichai.net with arXiv identifier 2605.30346 by researchers including You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, and Zhixiang Wang.
Paper Overview
Field: Computer Vision (CV)
Authors: You-Zhe Xie, Yu-Hsuan Li, Jie-Ying Lee, Kaipeng Zhang, Yu-Lun Liu, Zhixiang Wang
Posted: 2026-05-28
arXiv: 2605.30346
Summary
As video diffusion models (VDMs) advance toward becoming world models, a key question arises: do they truly understand causality, or are they merely overfitting to statistical temporal patterns? Existing benchmarks mostly rely on synthetic data, and the sim-to-real gap limits real-world generalization.
The authors propose YoCausal, a two-level benchmark inspired by the violation-of-expectation (VoE) paradigm from cognitive science. By time-reversing real-world videos at zero cost to serve as natural counterfactual samples, YoCausal establishes a freely scalable evaluation protocol.
- Level 1 — Reverse Surprise Index (RSI): quantifies a model's perception of the temporal arrow via denoising loss.
- Level 2 — Causal Cognition Index (CCI): uses vision-language models to stratify datasets into causal and non-causal subsets, disentangling genuine causal reasoning from temporal bias.
Findings
Evaluations across 13 state-of-the-art VDMs reveal that perceiving the arrow of time does not imply an understanding of causality, and a significant gap remains relative to human-level causal cognition.
---
*Auto-collected on 2026-06-01*
This page is an English static mirror generated for search and AI citation.
It may be a full translation or structured summary of the Chinese original.
Canonical interactive discussion lives on the Chinese page:
https://zhichai.net/topic/177980673