CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation
Mohamed bin Zayed University of Artificial Intelligence · Sailplane · Korea Advanced Institute of Science & Technology · Universidade Federal de Juiz de Fora and Centro Universitário Academia · Atal Bihari Vajpayee Indian Institute of Information Technology and Management, Gwalior · Department of Computer Science, Indian Institute of Technology, Madras, Indian Institute of Technology, Madras and National Institute of Information and Communications Technology (NICT), National Institute of Advanced Industrial Science and Technology · VTT Technical Research Centre of Finland Ltd. · Centre for Development of Telematics · Universidad Nacional de Córdoba · Technische Universität Darmstadt · NEC and Technische Universität Darmstadt · Copenhagen University · National Institute of Information and Communications Technology (NICT) · Federal University of Juiz de Fora · Mohamed bin Zayed University of Artificial Intelligence and University of Houston
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.findings-emnlp.1220 ↗
摘要
Translating cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meanings. In this work, we investigate whether images can act as cultural context in multimodal translation. We introduce CaMMT, a human-curated benchmark of over 5,800 triples of images along with parallel captions in English and regional languages. Using this dataset, we evaluate five Vision Language Models (VLMs) in text-only and text+image settings. Through automatic and human evaluations, we find that visual context generally improves translation quality, especially in handling Culturally-Specific Items (CSIs), disambiguation, and correct gender marking. By releasing CaMMT, our objective is to support broader efforts to build and evaluate multimodal translation systems that are better aligned with cultural nuance and regional variations.