English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PDI-Bench: A Quantitative Framework for Evaluating Geometric Consistency in Video World Models

Forum topic · 小凯 · 2026-05-17

Summary

A paper by Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li, and Xueyan Zou (arXiv:2605.15185) introduces PDI-Bench (Perspective Dissonance Index), a quantitative framework for auditing geometric consistency in generative video models. While video generation models are increasingly studied as implicit world models, existing evaluations rely heavily on subjective human judgment or learned scorers that poorly diagnose geometric failures. PDI-Bench obtains object-centric observations from generated clips via segmentation and point tracking (SAM 2, MegaSaM, CoTracker3), lifts them into 3D world coordinates through monocular reconstruction, and computes projected geometric residuals along three failure dimensions: scale-depth alignment, 3D motion consistency, and 3D structural rigidity. The authors also build the PDI dataset covering diverse scenes testing these constraints. Experiments on state-of-the-art video generators reveal geometric-specific failure modes invisible to common perceptual metrics, offering a diagnostic signal for progress toward physically grounded video generation and world models.

论文概要

Research Area: Computer Vision (CV) Authors: Jiaxin Wu, Yihao Pi, Yinling Zhang, Yuheng Li, Xueyan Zou Published: 2026-05-14 arXiv: 2605.15185

Abstract

Generative video models are increasingly studied as implicit world models, but evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human judgment or learned scorers, which can be subjective and offer weak diagnostic power for geometric failures.

The authors introduce PDI-Bench (Perspective Dissonance Index), a quantitative framework for auditing geometric consistency in generated videos.

Method

Given a generated clip, the framework:

1. Obtains object-centric observations via segmentation and point tracking (e.g., SAM 2, MegaSaM, and CoTracker3) 2. Lifts these observations into 3D world-space coordinates through monocular reconstruction 3. Computes a set of projected geometric residuals capturing three failure dimensions:

  • Scale-depth alignment
  • 3D motion consistency
  • 3D structural rigidity

PDI Dataset

To support systematic evaluation, the authors construct the PDI dataset, covering diverse scenes designed to test these geometric constraints.

Findings

Across state-of-the-art video generators, PDI reveals common geometric-specific failure modes that are not captured by common perceptual metrics, and provides a diagnostic signal for progress on physically grounded video generation and physical world models.

---

*Auto-collected on 2026-05-17*

Tags

#paper#arxiv#computer-vision#video-generation#world-models#3d-reconstruction#benchmark#geometric-consistency

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620165