← 返回论文检索
ACM Multimedia 2025Generative AI: Generative Multimedia

Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation

Weitao You, Heda Zuo, Junxian Wu 0003, Dengming Zhang, Zhibin Zhou 0002, Lingyun Sun

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755523 ↗

摘要

Achieving high-quality output alongside enhanced controllability is crucial in video-to-music generation, especially for optimizing user experience in real-life application scenarios. Most existing studies emphasize generative quality, but often overlooking the vital aspect of controllability. Therefore, the generated music cannot be easily fine-tuned or modified to meet users' expectations. In this paper, we delve into the spatial-temporal decomposition and alignment in controllable video-to-music generation. We first introduce a novel video-music decomposition and transformation approach in both spatial and temporal domain, and enhance the cross-modal correspondence through feature alignment and flow-matching based alignment. Furthermore, our method attains unsupervised controllability during training via feature-free guidance. Experimental results demonstrate that our model achieves state-of-the-art results in overall generative quality. Moreover, its controllability significantly outperforms existing models, making it exceptionally well-suited to accommodate users' flexible and diverse control requirements.