[论文] 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

研究领域: CV 作者: Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu 发布时间: 2026-10-02 arX…

目录
  1. 论文概要
  2. 中文摘要
  3. 原文摘要

论文概要

研究领域: CV 作者: Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu 发布时间: 2026-10-02 arXiv: 2610.03715

中文摘要

我们提出 4DCodeBench——一个通过代码生成进行四维逆图形的基准测试,要求智能体将视频中的动态场景重建为可执行的图形程序。为此,智能体必须将视觉观察转化为场景结构和动态的紧凑表示,通过实现物理仿真等抽象来复现复杂行为。为评估这一能力,我们整理了一组真实世界视频,并构建了涵盖形变、流体流动和断裂等多种物理现象的合成场景。我们对前沿模型进行了广泛基准测试,发现强大的静态重建能力尚不能转化为对复杂动态的可靠重建。4DCodeBench 为追踪'智能体通过代码理解世界动态'的进展提供了试验台。

原文摘要

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progre...


*自动采集于 2026-10-06*

#论文 #arXiv #CV #小凯

暂无表态

想参与讨论或点赞?登录后使用完整功能

讨论回复(0)

暂无回复,登录后可参与讨论

本文标签

合作

智谱 GLM-5 已上线

在智谱开放平台 BigModel.cn 打造 AI 应用。新一代旗舰模型 GLM-5 在推理、代码、智能体综合能力达到开源模型 SOTA。

领取 2000万 Tokens