Paper Overview
- Field: Computer Vision (CV)
- Authors: Hyeonwoo Kim, Jeonghwan Kim, Kyungwon Cho
- Published: 2026-04-22
- arXiv: 2604.20841
- Recent advances in video generative models enable synthesis of realistic human-object interaction (HOI) videos across broad scenarios and object categories, including complex dexterous manipulations that are hard to capture with motion capture systems.
- Despite the rich interaction knowledge in these synthetic videos, their limited physical fidelity and purely 2D nature make them difficult to use directly as imitation targets for physics-based character control.
- DeVI (Dexterous Video Imitation) leverages text-conditioned synthetic videos to enable physically plausible dexterous agent control for interacting with unseen target objects.
- To overcome the imprecision of generative 2D cues, the authors introduce a hybrid tracking reward that integrates 3D human tracking with robust 2D object tracking.
- Unlike prior approaches requiring high-quality 3D kinematic demonstrations, DeVI needs only generated videos, achieving zero-shot generalization across diverse objects and interaction types.
- Extensive experiments show DeVI outperforms existing methods that imitate 3D HOI demonstrations, especially in modeling dexterous hand-object interaction.
- The framework is further validated in multi-object scenarios and text-driven action diversity, highlighting the advantage of using video as an HOI-aware motion planner.
Key Points
Results
*Auto-collected on 2026-04-24.*