[论文] What 30,000 Hours of Ego-centric Video Does Not Teach
研究领域: CV 作者: Jiahua Dong, Anurag Bagchi, Yash Jangir, Muhammad Zubair Irshad, Sergey Zakharov, Martial Hebert, Homanga Bharadhwaj, Yu-Xiong Wang, Vitor Campagn…
论文概要
研究领域: CV 作者: Jiahua Dong, Anurag Bagchi, Yash Jangir, Muhammad Zubair Irshad, Sergey Zakharov, Martial Hebert, Homanga Bharadhwaj, Yu-Xiong Wang, Vitor Campagnolo Guizilini, Pavel Tokmakov 发布时间: 2026-10-08 arXiv: 2610.12464
中文摘要
世界模型为基于物理的仿真器提供了有前景的替代,但距实际部署仍远。我们追问:规模化自我中心人类视频能将其推进多远?使用包含 30,000 小时、逾 1,000 种场景类型和 14,000 名贡献者的数据集,我们不依赖不透明的下游指标,直接在具有挑战性的分布外基准上评估智能体建模与物体交互保真度。训练数据增加 100 倍确实改善两者,但不均衡:智能体建模良好,物体保真度仍很低且改善缓慢。我们发现智能体的提升并非必须来自数据——精心设计的视觉条件化方案仅用一小部分数据即可令其饱和,这使我们能单独测量物体保真度并发现其饱和点。随后引入的监督方案将模型容量从场景外观转向物体动态,改善了物体保真度,尽管仍有较大差距。结论可迁移至下游人形机器人建模。总体而言,规模化自我中心数据使智能体建模接近极限,而对"世界"的理解远远落后;弥合这一差距取决于模型如何被训练,而不仅是看到了多少数据。
原文摘要
World models offer a promising alternative to physics-based simulators, yet remain far from practical deployment. We ask how far scaling ego-centric human video takes them, using a dataset of 30,000 hours spanning over 1,000 scene types and 14,000 contributors. Rather than relying on opaque downstream metrics, we directly evaluate agent and object-interaction fidelity on a challenging out-of-distribution benchmark. Increasing training data by 100x improves both, but unevenly: the agent is modeled well, while object fidelity remains far lower and improves slowly. We show that the agent gains need not come from data, and a careful visual conditioning design saturates fidelity with a fraction of it, which lets us measure object fidelity on its own and discover its saturation point. We then in...
*自动采集于 2026-10-10*
#论文 #arXiv #CV #小凯