IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
Multimodal Art Projection · Kling Team, Kuaishou Technology · Hohai University · Beijing University of Posts and Telecommunications · Institute of automation, Chinese academy of science, Chinese Academy of Sciences · China University of Geoscience Beijing · Fudan University · Chongqing University · SUN YAT-SEN UNIVERSITY · University of Science and Technology Beijing · Guangdong OPPO Mobile Telecommunications Corp.,Ltd. · ByteDance Inc. · Alibaba Group · Mohamed bin Zayed University of Artificial Intelligence · Shenzhen University of Advanced Technology · Nanjing University · 01.AI · University of Waterloo · ByteDance Inc./TikTok
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Current benchmarks for Multimodal Large Language Models (MLLMs) predominantly rely on text-only queries, overlooking the essential role of images as visual context for enhancing video comprehension and facilitating natural human-AI interaction. To bridge this gap, we introduce \textbf{IV-Bench}, the first comprehensive benchmark for evaluating MLLMs on Image-Grounded Video Perception and Reasoning. IV-Bench comprises 966 videos paired with 2,560 meticulously annotated image-text queries across 13 tasks (7 perception and 6 reasoning tasks) spanning 5 distinct categories. We extensively evaluate state-of-the-art MLLMs, including open-source models (e.g., InternVL2.5, Qwen2.5-VL) and closed-source models (e.g., GPT-4o, Gemini2.0 series), revealing substantial performance gaps, with the best-performing model achieving only 28.9\% accuracy. Ablation studies demonstrate that incorporating images significantly enhances video understanding and highlight key model design factors influencing performance. Our findings provide valuable insights and guidance for future research. The code and dataset are available at \url{https://github.com/multimodal-art-projection/IV-Bench}.