RealVG: Unleashing MLLMs for Training-Free Spatio-Temporal Video Grounding in the Wild
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755381 ↗
摘要
Spatio-Temporal Video Grounding (STVG) aims to localize spatio-temporal tubes of specific objects or actions within videos based on textual queries. Despite significant progress, existing methods struggle to generalize effectively to real-world scenarios due to the limited quantity and diversity of annotated data. In this paper, we introduce RealVG, a robust and training-free pipeline that leverages powerful Multimodal Large Language Models (MLLMs) through question-answering to tackle STVG in the wild. To address the challenges posed by complex real-world videos and queries, we propose a spatio-temporal decoupling module and a query-guided visual token filter to decompose intricate scenes and refine target-oriented perception, enhancing the robustness and adaptability of MLLMs. Specifically, the spatio-temporal decoupling module breaks down videos and queries into simpler sub-scenes and sub-queries, reducing complexity and promoting a precise understanding of static visual elements. Meanwhile, the query-guided visual token filter eliminates irrelevant tokens, sharpening focus on the target object and improving short-range action perception. Experimental results demonstrate that RealVG achieves superior performance over state-of-the-art supervised and weakly supervised methods in real-world settings, despite requiring no STVG data for training.