← 返回论文检索
ACM Multimedia 2025Grand Challenges

MGVC: MLLM-Guided Video Captioning for the IntentVC Challenge

Zhipeng Yu, Qianqian Xu 0001, Yangbangyan Jiang, Pinci Yang, Qingming Huang

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3762058 ↗

摘要

Recently, with the rapid advancement of multimodal large language models (MLLMs), intent-oriented video captioning has received increasing attention due to its potential for controllable and grounded visual understanding. Fine-grained localized video captioning presents unique challenges due to the need for controllability, object grounding, and temporal precision. In this paper, we propose MGVC, a two-stage framework for intention-oriented controllable video captioning in the IntentVC 2025 Challenge. Our pipeline first leverages a fine-tuned MLLM to generate diverse preliminary captions. These candidate captions are then refined by another finetuned MLLM for further semantic alignment and stylistic coherence. We introduce a video-text matching module, further finetuned on the IntentVC dataset. This module will filter out semantically misaligned candidate captions. For caption selection, we train category-specific regressors that predict caption quality scores based on VTM similarity, textual features, intra-caption BLEU, and CLIP-based retrieval correlations. The caption with the highest predicted alignment score is chosen as final output. Finally, our method achieves 1st place in the IntentVC 2025 Grand Challenge, which demonstrates the effectiveness and generalization of our proposed method.