← 返回论文检索
ACM Multimedia 2025Datasets

HAN: Korean Heritage Augmented Narrative Visual-Language Description Dataset

SungHyun Moon, Aidyn Zhakatayev, SeungJae Lee

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3758229 ↗

摘要

The rapid advancements in Artificial Intelligence (AI) increasingly underscore the critical importance of multi-modal datasets for training robust and versatile models. However, the dominance of English-centric resources and the scarcity of datasets capturing diverse languages and cultural nuances limit AI's inclusivity and global applicability, posing a key challenge. Critically, language's intrinsic link to cultural context demands models possess profound cultural insight beyond mere translation to genuinely grasp societal norms and unique cultural expressions. Culturally rich broadcast content offers a solution; its systematic curation can mitigate linguistic/cultural imbalances, fostering culturally deep multi-modal datasets. This paper introduces the HAN (Korean Heritage Augmented Narrative Visual-Language Description) dataset, a new resource for multilingual image captioning and retrieval. HAN comprises 41,000 images captured from Korean broadcast video clips with 410,000 Korean/English narrative-style captions offering multifaceted perspectives on each visual instance. By incorporating Korean heritage, HAN reflects cultural diversity, addressing limitations of existing datasets focused on simple descriptions. Furthermore, this work analyzes HAN's caption diversity impact on retrieval, proposing strategies to enhance efficacy. These findings underscore HAN's potential to significantly advance multi-modal and multilingual processing, supported by its value as a rich resource for learning approaches that integrate diverse data types (like vision and text), natural language processing across various languages, and Korean heritage studies.