Deform360: A Massive Multi-View Visuotactile Dataset for Deformable Object World Modeling
Research area: Computer Vision (CV)
Authors: Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, and Yunzhu Li
Publication date: July 6, 2026
arXiv: 2607.05390
Abstract
Predicting object dynamics is a fundamental challenge in robotic manipulation, particularly for deformable objects. Their high-dimensional state spaces and complex material properties make accurate world modeling substantially more difficult than modeling rigid objects. Existing world models learn dynamics in either 2D pixel space or 3D geometric space, but the field has lacked large-scale real-world datasets for systematically comparing their respective strengths and limitations.
This paper presents Deform360, a large-scale visuotactile dataset containing 198 everyday objects, 1,980 interaction sequences, and more than 215 hours of observations. The data are collected using 41 cameras arranged around the manipulation area together with bimanual tactile grippers. This setup captures both global object motion and localized deformation caused by contact.
The authors introduce a markerless visual-tactile 3D tracking pipeline to extract dense geometric and motion information from the recorded interactions. They use the dataset to evaluate current state-of-the-art world models and compare 2D video-based models with 3D particle-based models.
Project website: https://deform360.lhy.xyz
Key points
- Focuses on world modeling for deformable-object robotic manipulation.
- Includes 198 everyday objects and 1,980 interaction sequences.
- Provides more than 215 hours of multi-view visuotactile observations.
- Combines 41 surrounding cameras with bimanual tactile grippers.
- Captures global motion and contact-induced local deformation.
- Uses a markerless visual-tactile 3D tracking workflow to obtain dense geometry and motion data.
- Systematically compares state-of-the-art 2D video and 3D particle world models.