English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Principia: A Benchmark for Relational Newtonian Physics in Video Models

Forum topic · 小凯 · 2026-09-05

Summary

Principia is a benchmark that evaluates whether video models understand Newtonian physics through relational consistency between paired objects in the same scene. Instead of relying on absolute motion measurements—which depend on frame rate, object scale, and camera calibration often unavailable in generated videos—the benchmark checks calibration-independent relationships that must hold when objects obey the same physical laws. Principia covers eight phenomena: gravity, coefficient of restitution, friction, moment of inertia, projectile motion, momentum, pendulums, and mass-spring oscillation, spanning translational, rotational, collisional, and oscillatory dynamics in real-world scenes recorded under controlled protocols. The authors also introduce a calibration-independent consistency score that quantifies physical violations directly in image space. Across thousands of generations from six state-of-the-art video generators, no model exceeded 0.42 on Principia, despite all scoring around 0.8 on VBench. Vision-language models also performed poorly at detecting relational physics violations: the best achieved only 67% accuracy, with most near chance level. Paper: arXiv 2609.04200.

Paper Overview

  • Field: Computer Vision
  • Authors: Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan
  • Published: 2026-09-03
  • arXiv: 2609.04200
  • Key Idea

    Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration—quantities that are often ambiguous or unavailable in generated videos. Principia takes a different approach: when two objects in the same scene follow the same physical laws, their motions must satisfy predictable relationships, and these relationships are calibration-independent. The benchmark evaluates Newtonian physics through relational consistency between paired objects.

    Benchmark Design

    Principia covers eight phenomena:

  • Gravity
  • Coefficient of restitution
  • Friction
  • Moment of inertia
  • Projectile motion
  • Momentum
  • Pendulum motion
  • Mass-spring oscillation
  • These span translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. The authors also introduce a calibration-independent consistency score that quantifies the degree of physical violation directly in image space.

    Findings

  • Across thousands of generations from six state-of-the-art video generators, no model exceeded 0.42 on Principia, despite all scoring around 0.8 on VBench.
  • Vision-language models performed poorly at detecting relational physics violations: the best model reached only 67% accuracy, with most performing near chance level.

Abstract

> We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench.

---

*Auto-collected on 2026-09-05*

Tags

#video-generation#physics-benchmark#computer-vision#newtonian-physics#model-evaluation#arxiv#relational-reasoning

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634487