← 返回论文检索
ICML 2026PosterAccept (regular)

OmniShow: Orchestrating Multimodal Conditions for Human-Object Interaction Video Generation

Donghao Zhou, Guisheng Liu, Hao Yang, Jiatong Li, Jingyu Lin, Xiaohu Huang, Yichen Liu, Xin Gao, Cunjian Chen, Shilei Wen, Chi Wing Fu, Pheng Ann Heng

The Chinese University of Hong Kong · ByteDance · Bytedance · ByteDance Inc. · Monash University · The University of Hong Kong · University of the Chinese Academy of Sciences · Miyou Network Technology (Shanghai) Co., Ltd. Hangzhou Branch · bytedance

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

In this work, we study **Human-Object Interaction Video Generation (HOIVG)**, which aims to synthesize high-quality HOI videos via text, reference image, audio, and pose conditions. To address the challenges of harmonious multimodal injection and heterogeneous data utility, we present **OmniShow**, an end-to-end framework tailored for HOIVG. We introduce *Unified Channel-wise Conditioning* to efficiently inject image and pose cues, *Gated Local-Context Attention* to ensure precise audio-visual synchronization, and a *Decoupled-then-Joint Training strategy* to effectively harness heterogeneous data. Extensive experiments on the proposed *HOIVG-Bench* demonstrate that OmniShow achieves state-of-the-art performance.