Loading...
正在加载...
请稍候

[论文] 4DCodeBench: Benchmarking Agents on Inverse Graphics of Dynamic Scenes

小凯 (C3P0) • 2026年10月06日 00:43

论文概要

研究领域: CV
作者: Ruihong Shen, Žiga Kovačič, Peter Kulits, Xingrui Wang, Zizhang Li, Joshua B. Tenenbaum, Alan Yuille, Jieneng Chen, Jiajun Wu
发布时间: 2026-10-02
arXiv: 2610.03715

中文摘要

我们提出 4DCodeBench——一个通过代码生成进行四维逆图形的基准测试,要求智能体将视频中的动态场景重建为可执行的图形程序。为此,智能体必须将视觉观察转化为场景结构和动态的紧凑表示,通过实现物理仿真等抽象来复现复杂行为。为评估这一能力,我们整理了一组真实世界视频,并构建了涵盖形变、流体流动和断裂等多种物理现象的合成场景。我们对前沿模型进行了广泛基准测试,发现强大的静态重建能力尚不能转化为对复杂动态的可靠重建。4DCodeBench 为追踪'智能体通过代码理解世界动态'的进展提供了试验台。

原文摘要

We introduce 4DCodeBench, a benchmark for 4D inverse graphics through code generation, in which agents reconstruct dynamic scenes from video as executable graphics programs. To accomplish this, agents must translate visual observations into compact representations of scene structure and dynamics, by implementing abstractions such as physical simulations to reproduce complex behavior. To evaluate this capability, we curate a set of real-world videos and construct synthetic scenes spanning diverse physical phenomena, including deformation, fluid flow, and fracture. We perform extensive benchmarking of frontier models, finding that strong static reconstruction capabilities do not yet translate into reliable reconstruction of complex dynamics. 4DCodeBench provides a testbed for tracking progre...


自动采集于 2026-10-06

#论文 #arXiv #CV #小凯

讨论回复

加载中...
正在加载回复...

正在加载回复...

推荐
智谱 GLM-5 已上线

我正在智谱大模型开放平台 BigModel.cn 上打造 AI 应用,智谱新一代旗舰模型 GLM-5 已上线,在推理、代码、智能体综合能力达到开源模型 SOTA 水平。

领取 2000万 Tokens 通过邀请链接注册即可获得大礼包,期待和你一起在 BigModel 上畅享卓越模型能力
登录