← 返回论文检索
ACM Multimedia 2025Generative AI: Social Aspects of Generative AI

Towards Culturally Fair Multimodal Generation: Quantifying and Mitigating Orientalist Biases in Text-to-Visual Models

Yifan Zeng, Fangzhou Dong, Jian Zhao 0013, Peijia Zheng, Jian Li 0034, Huiyu Zhou 0005

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755251 ↗

摘要

This study systematically uncovers and quantitatively evaluates the pervasive Orientalist biases in text-to-image (T2I) and text-to-video (T2V) generation models through a sociocultural lens grounded in postcolonial Orientalist theoretical frameworks. We identify systematic biases in the visual representations produced by multimodal generative models, including hyper-exoticization and temporal alienation. These biases mirror colonial-era narratives and undermine equitable sociocultural communication. Through empirical analysis of 8 mainstream T2I models and 4 T2V models, we demonstrate that culturally neutral prompts related to China consistently generate visual outputs embedded with Orientalist biases. We develop a novel visual question answering (VQA) framework as an evaluation metric, leveraging state-of-the-art vision-language model (VLM) to establish the first automated quantitative assessment methodology for such biases. A mitigation framework employing large language model (LLM) is proposed and experimentally validated. This interdisciplinary work illuminates the societal implications of multimodal generative models while advancing efforts toward fair and inclusive social computing.