Crowdsource, Crawl, or Generate? Creating SEA-VL, a Multicultural Vision-Language Dataset for Southeast Asia
Cohere · SEACrowd · DoorDash · Independent Researcher · Stanford University · Universitas Indonesia · Oracle · University of Bath · Samsung · Capital One · University of Central Florida and Universitas Islam Indonesia · Samsung Research · Meta · SCB 10X · Sony · Hanyang University · Tianjin University · Binus University · Institut Teknologi Sepuluh Nopember · University of Central Florida and Universitas Gadjah Mada · National University of Singapore · King Mongkut’s University of Technology Thonburi · Institut Teknologi Bandung · Technical University of Wroclaw · Independent · Nanyang Technological University · Modelcode.ai · Vidyasirimedhi Institute of Science and Technology · Insignia · Mohamed bin Zayed University of Artificial Intelligence · Indian Statistical Institute · , A*STAR · Singapore Polytechnic · University of the Philippines · Carnegie Mellon University · Dataxet:Sonar · , University of British Columbia · Allen Institute for Artificial Intelligence · Universitas Brawijaya · Pelita Harapan University · Singapore University of Technology and Design · Brown University · Beijing Academy of Artificial Intelligence (BAAI) · AI Singapore
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.acl-long.916 ↗
摘要
Despite Southeast Asia’s (SEA) extraordinary linguistic and cultural diversity, the region remains significantly underrepresented in vision-language (VL) research, resulting in AI models that inadequately capture SEA cultural nuances. To fill this gap, we present SEA-VL, an open-source initiative dedicated to developing culturally relevant high-quality datasets for SEA languages. By involving contributors from SEA countries, SEA-VL ensures better cultural relevance and diversity, fostering greater inclusivity of underrepresented languages and cultural depictions in VL research. Our methodology employed three approaches: community-driven crowdsourcing with SEA contributors, automated image crawling, and synthetic image generation. We evaluated each method’s effectiveness in capturing cultural relevance. We found that image crawling achieves approximately ~85% cultural relevance while being more cost- and time-efficient than crowdsourcing, whereas synthetic image generation failed to accurately reflect SEA cultural nuances and contexts. Collectively, we gathered 1.28 million SEA culturally relevant images, more than 50 times larger than other existing datasets. This work bridges the representation gap in SEA, establishes a foundation for developing culturally aware AI systems for this region, and provides a replicable framework for addressing representation gaps in other underrepresented regions.