OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation
The Chinese University of Hong Kong · ByteDance · Bytedance · ByteDance Inc. · Monash University · The University of Hong Kong · University of the Chinese Academy of Sciences · Miyou Network Technology (Shanghai) Co., Ltd. Hangzhou Branch · bytedance
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
In this work, we study **Human-Object Interaction Video Generation (HOIVG)**, which aims to synthesize high-quality HOI videos via text, reference image, audio, and pose conditions. To address the challenges of harmonious multimodal injection and heterogeneous data utility, we present **OmniShow**, an end-to-end framework tailored for HOIVG. We introduce *Unified Channel-wise Conditioning* to efficiently inject image and pose cues, *Gated Local-Context Attention* to ensure precise audio-visual synchronization, and a *Decoupled-then-Joint Training strategy* to effectively harness heterogeneous data. Extensive experiments on the proposed *HOIVG-Bench* demonstrate that OmniShow achieves state-of-the-art performance.