A Very Big Video Reasoning Suite
University of California, Berkeley · Sensetime · Northeastern University · University of California, San Diego · Max Planck Institute for Intelligent Systems · Johns Hopkins University · University of Michigan - Ann Arbor · Amazon · Epsilon · Facebook · Shanghai Jiao Tong University · East China Normal University · Stanford University · University of Texas at Austin · Amazon Web Services · Independent Researcher · Nanyang Technological University · AGI, Amazon · ByteDance Inc. · Tesla · University of California, Irvine · Technical University of Munich · Imperial College London · University of Edinburgh, University of Edinburgh · The Hong Kong University of Science and Technology (Guangzhou) · Auburn University · Columbia University · University College London · University of Michigan · University of Oxford · Vector Institute · UCSD, USA · Harvard · Tsinghua University · The Chinese University of Hong Kong · Univ. Of Washington · Sensetime Ltd.
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality. Systematically studying video reasoning and its scaling behavior suffers from a lack of video reasoning (training) data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks and over one million video clips—approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, toolkit, and models will be released publicly.