← 返回论文检索
ACM Multimedia 2024Technical Demonstrations

MAF-ID: Multi-Agent Framework for Interactive Dubbing through Deep Video Understanding

Zhanbin Hu, Xiaodong He, Renzhou Pan, Xianzhou Zeng, Chenming Fan, Qiang Zhu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3664647.3684992 ↗

摘要

In the domain of video generation, Text-to-video suffers from a notable application gap due to lack of audio that harmonizes with the visual content. Current solutions typically dubbing based solely on the original text used for generate video, which causes a mismatch between the video content and audio details, primarily stems from the lack of understanding of the video's visual modality. Leveraging advancements in multimodal large language model and LLM-based Agent, we propose MAF-ID, a multi-agent interactive framework for video dubbing based on deep video understanding. MAF-ID achieves agent collaboration through the autonomous interaction of three agents, to capture a deep understanding of the video visual content from macro to micro, progressively generate sound effects, voice-overs, and background music that is adaptive to the video. By deeply aligning text, video, and audio modalities, our method significantly enhances the fine-grained coordination between video and audio, making it widely available for AI-generated videos, VLOGs, and other video production scenarios requiring dubbing.