Afri-MCQA: Multimodal Cultural Question Answering for African Languages
Department of Computer Science, University College London · Lesan AI · Friedrich-Alexander-Universität Erlangen-Nürnberg, ML Collective and Masakhane · University of Pretoria · Digital Umuganda · Mohamed bin Zayed University of Artificial Intelligence · McGill University · Mohamed bin Zayed University of Artificial Intelligence and University of Houston
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.1869 ↗
摘要
Africa is home to over one-third of the world’s languages, yet remains severely underrepresented in multimodal AI research. We introduce Afri-MCQA, the first Multilingual Cultural Question-Answering benchmark containing 7.5k Q A pairs across 15 African languages from 12 countries. The benchmark offers parallel text and speech modalities and was entirely created by native speakers. We find that models show poor performance across evaluated cultures, with near-zero accuracy on open-ended VQA when queried through native language or speech. To test linguistic competence, we include control experiments meant to assess this specific aspect separate from cultural knowledge, and we observe significant performance gaps between native languages and English for both text and speech. These findings underscore the pressing need for speech-first approaches, culturally grounded pretraining, and cross-lingual cultural transfer. We release Afri-MCQA to support more inclusive multimodal AI development.