English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

PDI-Bench: Quantitative Evaluation of Geometric Consistency in Video World Models

Forum topic · 小凯 · 2026-05-16

Summary

PDI-Bench (Perspective Distortion Index) is a quantitative framework for auditing geometric coherence in generative video models, proposed by Jiaxin Wu, Yihao Pi, and Yinling Zhang (arXiv:2505.08633). While generative video models are increasingly viewed as implicit world models, existing evaluation pipelines rely on subjective human judgment or learned graders that poorly diagnose geometric failures. PDI-Bench extracts object-centric observations from generated clips using segmentation and point tracking tools such as SAM 2, MegaSaM, and CoTracker3, lifts them to 3D world-space coordinates via monocular reconstruction, and computes projective-geometry residuals across three failure dimensions: scale-depth alignment, 3D motion consistency, and 3D structural rigidity. The authors also release the PDI dataset covering diverse scenes designed to test these geometric constraints, along with code at https://pdi-bench.github.io/. Experiments on state-of-the-art video generators reveal consistent geometry-specific failure modes missed by perceptual metrics, providing a diagnostic signal toward physically grounded video generation and world models.

Paper Overview

  • Research area: Computer Vision (CV)
  • Authors: Jiaxin Wu, Yihao Pi, Yinling Zhang
  • Published: 2026-05-16
  • arXiv: 2505.08633
  • Summary

    Generative video models are increasingly studied as implicit world models, yet evaluating whether they produce physically plausible 3D structure and motion remains challenging. Most existing video evaluation pipelines rely heavily on human judgment or learned graders, which can be subjective and weakly diagnostic for geometric failures.

    The authors introduce PDI-Bench (Perspective Distortion Index), a quantitative framework for auditing geometric coherence in generated videos:

  • Given a generated clip, object-centric observations are obtained via segmentation and point tracking (e.g., SAM 2, MegaSaM, and CoTracker3).
  • Observations are lifted to 3D world-space coordinates via monocular reconstruction.
  • A set of projective-geometry residuals is computed, capturing three failure dimensions:
  • 1. Scale-depth alignment 2. 3D motion consistency 3. 3D structural rigidity

    Dataset and Findings

    To support systematic evaluation, the authors construct the PDI dataset, covering diverse scenes designed to test these geometric constraints. Across state-of-the-art video generators, PDI reveals consistent geometry-specific failure modes not captured by common perceptual metrics, offering a diagnostic signal for progress toward physically grounded video generation and physical world models.

    Resources

  • Project page, code, and dataset: https://pdi-bench.github.io/
  • arXiv: https://arxiv.org/abs/2505.08633

Tags

#video-generation#world-models#computer-vision#evaluation-benchmark#3d-reconstruction#geometric-consistency#pdi-bench#arxiv

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177620090