Paper Overview
- Field: Computer Vision
- Authors: Varun Varma Thozhiyoor, Shivam Tripathi, Venkatesh Babu Radhakrishnan
- Published: 2026-09-03
- arXiv: 2609.04200
- Gravity
- Coefficient of restitution
- Friction
- Moment of inertia
- Projectile motion
- Momentum
- Pendulum motion
- Mass-spring oscillation
- Across thousands of generations from six state-of-the-art video generators, no model exceeded 0.42 on Principia, despite all scoring around 0.8 on VBench.
- Vision-language models performed poorly at detecting relational physics violations: the best model reached only 67% accuracy, with most performing near chance level.
Key Idea
Evaluating physical reasoning in video models is difficult because absolute motion measurements depend on frame rate, object scale, and camera calibration—quantities that are often ambiguous or unavailable in generated videos. Principia takes a different approach: when two objects in the same scene follow the same physical laws, their motions must satisfy predictable relationships, and these relationships are calibration-independent. The benchmark evaluates Newtonian physics through relational consistency between paired objects.
Benchmark Design
Principia covers eight phenomena:
These span translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. The authors also introduce a calibration-independent consistency score that quantifies the degree of physical violation directly in image space.
Findings
Abstract
> We introduce Principia, a benchmark that evaluates Newtonian physics through relational consistency between paired objects. Principia spans eight phenomena across translational, rotational, collisional, and oscillatory dynamics, using real-world scenes recorded under controlled protocols. Across thousands of generations from six state-of-the-art video generators, no model exceeds 0.42 on Principia despite all scoring around 0.8 on VBench.
---
*Auto-collected on 2026-09-05*