OmniDenseCap: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions
Peking University · South China University of Technology · University of Electronic Science and Technology of China · University of Hong Kong · Institute of automation, Chinese academy of science, Chinese Academy of Sciences · The University of Hong Kong · HKU, Reka · Kuaishou Technology · Alibaba Group · Kuaishou- 快手科技
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
This paper proposes Omni Dense Captioning, a novel task designed to generate continuous, fine-grained, and structured audio-visual narratives with explicit timestamps. To ensure dense semantic coverage, we introduce a six-dimensional structural schema to create "script-like" captions, enabling readers to vividly imagine the video content scene-by-scene, akin to a cinematographic screenplay. To facilitate research, we construct OmniDCBench, a high-quality human-annotated benchmark, and propose SodaM, a unified metric that evaluates time-aware detailed descriptions while mitigating scene boundary ambiguity. Furthermore, we construct a training dataset OmniDenseCap-40K and present Omni-Captioner-7B, a strong baseline trained via SFT and GRPO with task-specific rewards. Extensive experiments demonstrate that Omni-Captioner-7B achieves state-of-the-art performance, surpassing Gemini-2.5-Pro, while its generated dense descriptions significantly boost downstream capabilities in audio-visual reasoning (DailyOmni and WorldSense) and temporal grounding (Charades-STA). All datasets, models, and code will be made publicly available.