ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
Multimodal Art Projection · Hohai University · Shanghai University of Electric Power · University of Luxemburg · University of Science and Technology of China · Beijing Institute of Technology · Hong Kong University of Science and Technology · Beijing University of Posts and Telecommunications · Hubei Business College · China University of Geoscience Beijing · Mohamed bin Zayed University of Artificial Intelligence · Beihang University · University of Manchester · Guangdong OPPO Mobile Telecommunications Corp.,Ltd. · Nanjing University · 01.AI · University of Waterloo · Shenzhen University of Advanced Technology · ByteDance Inc./TikTok
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Although long-video understanding demands that models capture hierarchical temporal information—from clip and shot to event and story—existing benchmarks either neglect this multi-scale design or scatter scale-specific questions across different videos, preventing direct comparison of model performance across timescales on the same content. To address this, we introduce ScaleLong, the first benchmark to disentangle these factors by embedding questions targeting four hierarchical timescales\textemdash clip, shot, event, and story\textemdash all within the same video content. This within-content multi-timescale questioning design enables direct comparison of model performance across timescales on identical videos. ScaleLong features 269 long videos (avg. 86 min) from 5 main categories and 36 sub-categories, with 4–8 carefully designed questions, with at least one question targeting each timescale. Evaluating 23 MLLMs reveals a distinct U-shaped performance trend: higher accuracy at the shortest (clip) and longest (story) timescales, with a dip at intermediate levels. Furthermore, ablation studies demonstrate that increased visual token capacity consistently enhances reasoning across all timescales. ScaleLong offers a crucial fine-grained, multi-timescale benchmark for advancing MLLM capabilities in long-video understanding. The code and dataset are available at \url{https://github.com/multimodal-art-projection/ScaleLong}