English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

Forum topic · 小凯 · 2026-04-15

Summary

OmniShow is an end-to-end framework for Human-Object Interaction Video Generation (HOIVG), synthesizing high-quality videos conditioned on text, reference images, audio, and pose. Presented by researchers including Donghao Zhou and Chi-Wing Fu (arXiv:2604.11804), the work targets real-world applications such as e-commerce demonstrations, short video production, and interactive entertainment. The framework introduces unified channel conditioning for efficient injection of reference images and pose signals, and a gated local context attention mechanism to achieve precise audio-video synchronization. To address training data scarcity, the authors develop a disentangle-then-joint training strategy. They also contribute HOIVG-Bench, a dedicated benchmark for evaluating human-object interaction video generation. According to the paper's abstract, OmniShow delivers industrial-grade performance in coordinating multiple modality conditions, making it a notable reference for controllable, multimodal video synthesis research.

[Paper] OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

Paper Overview

  • Field: cs.CV
  • Authors: Donghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li, Jingyu Lin, Xiaohu Huang, Yichen Liu, Xin Gao, Cunjian Chen, Shilei Wen, Chi-Wing Fu, Pheng-Ann Heng
  • Published: 2026-04-13
  • arXiv: 2604.11804
  • Abstract (from the paper)

    In this work, we study Human-Object Interaction Video Generation (HOIVG), which aims to synthesize high-quality human-object interaction videos conditioned on text, reference images, audio, and pose. This task holds significant practical value for automating content creation in real-world applications, such as e-commerce demonstrations, short video production, and interactive entertainment.

    Key Contributions

  • OmniShow framework: an end-to-end model that coordinates multiple multimodal conditions (text, reference image, audio, pose) and delivers industrial-grade performance.
  • Unified channel conditioning: enables efficient injection of reference images and pose signals.
  • Gated local context attention: ensures precise audio-video synchronization.
  • Disentangle-then-joint training strategy: developed to mitigate data scarcity for this task.
  • HOIVG-Bench: a dedicated benchmark for human-object interaction video generation.

Applications

E-commerce product demonstrations, short-form video production, and interactive entertainment.

--- *Automatically collected on 2026-04-15.*

Tags

#video-generation#multimodal#human-object-interaction#diffusion-models#arxiv#computer-vision#ai-research

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/177618472