[Paper] OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation
Paper Overview
- Field: cs.CV
- Authors: Donghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li, Jingyu Lin, Xiaohu Huang, Yichen Liu, Xin Gao, Cunjian Chen, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng
- Published: 2026-04-13
- arXiv: 2604.11804
- OmniShow framework: an end-to-end model that coordinates multiple multimodal conditions (text, reference image, audio, pose) and delivers industrial-grade performance.
- Unified channel conditioning: enables efficient injection of reference images and pose signals.
- Gated local context attention: ensures precise audio-video synchronization.
- Disentangle-then-joint training strategy: developed to mitigate data scarcity for this task.
- HOIVG-Bench: a dedicated benchmark for human-object interaction video generation.
Abstract (from the paper)
In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditioned on text, reference images, audio, and pose. This task holds significant practical value for automating content creation in real-world applications, such as e-commerce demonstrations, short video production, and interactive entertainment.
Key Contributions
Applications
E-commerce product demonstrations, short-form video production, and interactive entertainment.
--- *Automatically collected on 2026-04-15.*