Query-Focused Multimodal Summarization with Gate-Guided Mixture-of-Experts
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754699 ↗
摘要
The goal of generic multimodal summarization is to extract the most important information from different modalities to form summaries. Yet the importance of scenes and text in a video is often subjective, and users should have the option of customizing the summary by using natural language to specify what is important to them. However, existing methods for fully automatic multimodal summarization have not exploited available language models, which can serve as an effective prior for saliency. To address this issue, we introduce Query-Focused Multimodal Summ arization(QFSumm), a single framework for addressing both generic and query-focused multimodal summarization, typically approached separately in the literature. In addition, we propose a novel gate-guided mixture-of-experts that uses expert gate module to organize three experts (video expert, text expert and shared expert) to model the correlations between multimodal information. In addition, we propose two novel contrastive losses to represent consistency and diversity. Extensive experiments on a query-focused video summarization dataset (QFVS), two standard video summarization datasets (TVSum and SumMe) and three multimodal summarization datasets (CNN, Daily Mail and BLiSS) demonstrate the superiority of QFSumm, achieving state-of-the-art performances on all datasets.