← 返回论文检索
ACM Multimedia 2025Datasets

Valor32k-AVQA v2.0: Open-Ended Audio-Visual Question Answering Dataset and Benchmark

Ines Riahi, Abduljalil Radman, Zixin Guo, Rachid Hedjam, Jorma Laaksonen

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3758261 ↗

摘要

Despite growing interest in Audio-Visual Question Answering (AVQA), existing datasets often suffer from limited diversity, rigid formats, and insufficient integration of audio and visual modalities. To address these limitations, we introduce Valor32k-AVQA v2.0, a large-scale dataset containing 28,863 real-world videos and over 225,000 QA pairs, designed to support diverse and realistic multimodal understanding. The dataset features both open-ended and multiple-choice questions, each annotated with the required modality ( visual, audio, or audio-visual ) and question category ( description, action, count, temporal, location, or relative position ). All annotations-including questions, answers, and metadata-are generated through a fully automated prompting pipeline using GPT-4o, with human validation performed on a representative sample to ensure quality. We benchmark a few state-of-the-art models, with additional evaluations available on the project page, and observe that incorporating audio consistently improves performance during fine-tuning without compromising visual reasoning capabilities. These findings highlight that the audio signals in our dataset are not only well integrated, but also informative and complementary, establishing Valor32k-AVQA v2.0 as a valuable resource for developing and evaluating robust audio-visual question answering systems.