← 返回论文检索
NeurIPS 2024Oral PosterAccept (Oral)

CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

David Romero, Chenyang Lyu, Haryo Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesus-German Ortiz-Barajas, Emilio Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Yong Zheng-Xin, Zheng Wei Lim, Paula Silva, Jocelyn Dunstan, Mélanie Jouitteau, David LE MEUR, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Gráinne Caulfield, Teresa Lynn, Christian Salamea-Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Azime, Henok Ademtew, Bontu Balcha, Naome A. Etori, David Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Imam, Kumaranage Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Niyomugisha Olivier, Jay Gala, Pranjal Chitale, Fauzan Farooqui, Thamar Solorio, Alham Aji

MBZUAI · Mohamed bin Zayed University of Artificial Intelligence · Universidad de la República · Technische Universität Darmstadt · Institute for Computer Science, Artificial Intelligence and Technology · KAIST · Korea Advanced Institute of Science & Technology · Mohamed bin Zayed University of Artificial Intelligence (MBZUAI) · Brown University · Google · Universidad de Santiago de Chile · University of Chile · CNRS · University of Michigan - Ann Arbor · United Arab Emirates University · Mongolian University of Science and Technology · Emirates Center for Mobility Research, UAE University · Universidad Nacional de Córdoba · Universidad Nacional de Córdoba, Argentina · Federal University of Juiz de Fora · Universidade Federal de Juiz de Fora · SAMSUNG ELECTRONICS PHILIPPINES CORPORATION · Santa Clara University · Skyline High School · University of Cambridge · University of Michigan · Dublin City University · Universidad Politécnica Salesiana · Pontificia Universidad Católica de Chile · Universität des Saarlandes · Ethiopian AI Institute · Addis Ababa Institute of Technology · University of Minnesota - Twin Cities · Mila & McGill University · Instituto Politécnico Nacional · Universität Stuttgart · University of Melbourne · AI Singapore · Universidad Politécnica de Madrid · Institut Teknologi Bandung · Cohere · National Institute of Information and Communications Technology (NICT) · Indian Institute of Technology, Madras, Dhirubhai Ambani Institute Of Information and Communication Technology · National Institute of Information and Communications Technology (NICT), National Institute of Advanced Industrial Science and Technology · State University of New York at Stony Brook · UdelaR · University of Rwanda · University of Mumbai · Microsoft Research · Visvesvaraya National Institute of Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Visual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field.