← 返回论文检索
ACM Multimedia 2024Oral Session 10: Speech and Audio in Multimedia Processing

TiVA: Time-Aligned Video-to-Audio Generation

Xihua Wang 0002, Yuyue Wang 0003, Yihan Wu 0008, Ruihua Song, Xu Tan 0003, Zehua Chen 0005, Hongteng Xu, Guodong Sui

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3664647.3681027 ↗

摘要

Video-to-audio generation is crucial for autonomous video editing and post-processing, which aims to generate high-quality audio for silent videos with semantic similarity and temporal synchronization. However, most existing methods mainly focus on matching the semantics of the visual and acoustic modalities while merely considering their temporal alignment in a coarse granularity, thus failing to achieve precise synchronization. In this study, we propose a novel time-aligned video-to-audio framework, called TiVA, to achieve semantic matching and temporal synchronization jointly when generating audio. Given a silent video, our method encodes its visual semantics and predicts an audio layout separately. Then, leveraging the semantic latent embeddings and the predicted audio layout as condition, it learns a latent diffusion-based audio generator. Comprehensive objective and subjective experiments demonstrate that our method consistently outperforms state-of-the-art methods on semantic matching and temporal synchronization.