← 返回论文检索
ICML 2026PosterAccept (regular)

A Very Big Video Reasoning Suite

Maijunxian Wang, Ruisi Wang, Juyi Lin, Ran Ji, Thaddäus Wiedemer, Qingying Gao, Dezhi Luo, Yaoyao Qian, Lianyu Huang, Zelong Hong, Jiahui Ge, Qianli Ma, Hang He, Yifan Zhou, Lingzi Guo, Lantao Mei, Jiachen Li, Hanwen Xing, Tianqi Zhao, Fengyuan Yu, Weihang Xiao, Yizheng Jiao, Jianheng Hou, Danyang Zhang, Pengcheng Xu, Boyang ZHONG, Zehong Zhao, Gaoyun Fang, John Kitaoka, Xu Yile, Hua XU, Kenton Blacutt, Tin Nguyen, Siyuan Song, Haoran Sun, shaoyue wen, Linyang He, Runming Wang, Yanzhi Wang, Mengyue Yang, Ziqiao Ma, Raphaël Millière, Freda Shi, Nuno Vasconcelos, Daniel Khashabi, Alan Yuille, Yilun Du, Ziming Liu, Dahua Lin, Ziwei Liu, Vikash Kumar, Yijiang Li, Lei Yang, Zhongang Cai, Hokin Deng

University of California, Berkeley · Sensetime · Northeastern University · University of California, San Diego · Max Planck Institute for Intelligent Systems · Johns Hopkins University · University of Michigan - Ann Arbor · Amazon · Epsilon · Facebook · Shanghai Jiao Tong University · East China Normal University · Stanford University · University of Texas at Austin · Amazon Web Services · Independent Researcher · Nanyang Technological University · AGI, Amazon · ByteDance Inc. · Tesla · University of California, Irvine · Technical University of Munich · Imperial College London · University of Edinburgh, University of Edinburgh · The Hong Kong University of Science and Technology (Guangzhou) · Auburn University · Columbia University · University College London · University of Michigan · University of Oxford · Vector Institute · UCSD, USA · Harvard · Tsinghua University · The Chinese University of Hong Kong · Univ. Of Washington · Sensetime Ltd.

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Video reasoning grounds intelligence in spatiotemporally consistent visual environments that go beyond what text can naturally capture, enabling intuitive reasoning over motion, interaction, and causality. Rapid progress in video models has focused primarily on visual quality. Systematically studying video reasoning and its scaling behavior suffers from a lack of video reasoning (training) data. To address this gap, we introduce the Very Big Video Reasoning (VBVR) Dataset, an unprecedentedly large-scale resource spanning 200 curated reasoning tasks and over one million video clips—approximately three orders of magnitude larger than existing datasets. We further present VBVR-Bench, a verifiable evaluation framework that moves beyond model-based judging by incorporating rule-based, human-aligned scorers, enabling reproducible and interpretable diagnosis of video reasoning capabilities. Leveraging the VBVR suite, we conduct one of the first large-scale scaling studies of video reasoning and observe early signs of emergent generalization to unseen reasoning tasks. Together, VBVR lays a foundation for the next stage of research in generalizable video reasoning. The data, toolkit, and models will be released publicly.