Overview
Field: Computer Vision / Robotics Authors: Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, Yunzhu Li arXiv: 2607.05390 Project site: https://deform360.lhy.xyz
Key points
- Predicting object dynamics (world modeling) is a fundamental challenge in robot manipulation. Deformable objects are particularly hard to model because of their high-dimensional state spaces and complex material properties.
- Current world models learn dynamics either in 2D pixel space or in 3D geometric space, but there is a lack of large-scale real-world data for systematically understanding the strengths and weaknesses of each approach.
- The paper introduces Deform360, a large-scale visuotactile dataset featuring:
- 198 everyday objects
- 1,980 interaction sequences
- 215+ hours of observation data
- Data capture setup:
- 41 surrounding-view cameras capturing global object motion
- Bimanual tactile grippers capturing contact-induced local deformations
- A markerless visuotactile 3D tracking pipeline is used to extract dense geometry and motion from the recordings.
- The authors benchmark state-of-the-art world models on the dataset, comparing 2D video-based models against 3D particle-based models.
- Paper: https://arxiv.org/abs/2607.05390
- Project website: https://deform360.lhy.xyz