English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

Deform360: A Multi-View Visuotactile Dataset for Deformable Object World Modeling

Forum topic · 小凯 · 2026-07-08

Summary

This paper introduces Deform360, a large-scale visuotactile dataset designed to advance world modeling for robotic manipulation of deformable objects. The dataset covers 198 everyday objects, 1,980 interaction sequences, and more than 215 hours of observations. It combines 41 surrounding cameras with bimanual tactile grippers, capturing both global object motion and localized deformation caused by physical contact. A markerless visual-tactile 3D tracking pipeline extracts dense geometry and motion information from the recorded interactions. The authors use this dataset to systematically evaluate state-of-the-art world models and compare 2D video-based approaches with 3D particle-based representations. Deform360 addresses the lack of large-scale real-world benchmarks for modeling deformable-object dynamics, which involve high-dimensional state spaces and complex material behavior. By integrating visual and tactile signals with multi-view observation, the dataset supports research on dynamics prediction, representation learning, and robotic manipulation. The project website is https://deform360.lhy.xyz, and the paper is available at https://arxiv.org/abs/2607.05390.

Deform360: A Massive Multi-View Visuotactile Dataset for Deformable Object World Modeling

Research area: Computer Vision (CV)

Authors: Hongyu Li, Wanjia Fu, Xiaoyan Cong, Zekun Li, Binghao Huang, Hanxiao Jiang, Xintong He, Yiqing Liang, Rao Fu, Tao Lu, Srinath Sridhar, Kevin A. Smith, George Konidaris, and Yunzhu Li

Publication date: July 6, 2026

arXiv: 2607.05390

Abstract

Predicting object dynamics is a fundamental challenge in robotic manipulation, particularly for deformable objects. Their high-dimensional state spaces and complex material properties make accurate world modeling substantially more difficult than modeling rigid objects. Existing world models learn dynamics in either 2D pixel space or 3D geometric space, but the field has lacked large-scale real-world datasets for systematically comparing their respective strengths and limitations.

This paper presents Deform360, a large-scale visuotactile dataset containing 198 everyday objects, 1,980 interaction sequences, and more than 215 hours of observations. The data are collected using 41 cameras arranged around the manipulation area together with bimanual tactile grippers. This setup captures both global object motion and localized deformation caused by contact.

The authors introduce a markerless visual-tactile 3D tracking pipeline to extract dense geometric and motion information from the recorded interactions. They use the dataset to evaluate current state-of-the-art world models and compare 2D video-based models with 3D particle-based models.

Project website: https://deform360.lhy.xyz

Key points

  • Focuses on world modeling for deformable-object robotic manipulation.
  • Includes 198 everyday objects and 1,980 interaction sequences.
  • Provides more than 215 hours of multi-view visuotactile observations.
  • Combines 41 surrounding cameras with bimanual tactile grippers.
  • Captures global motion and contact-induced local deformation.
  • Uses a markerless visual-tactile 3D tracking workflow to obtain dense geometry and motion data.
  • Systematically compares state-of-the-art 2D video and 3D particle world models.
*Source metadata was automatically collected on July 6, 2026.*

Tags

#deformable-objects#robotics#world-models#visuotactile-dataset#multi-view-learning#computer-vision

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178346217