VisTW: Benchmarking Vision-Language Models for Taiwanese Mandarin in Taiwan
National Taiwan University and Appier · National Taiwan University · University of Illinois at Urbana-Champaign
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.1830 ↗
摘要
Vision-Language Models (VLMs) often struggle in Taiwanese Mandarin environments due to region-specific orthographic and cultural context. We introduce VisTW, a comprehensive benchmark featuring (i) multiple-choice questions (3,795 academic questions) and (ii) free-form generation evaluation (141 Taiwanese-context free-form pairs). Beyond standard accuracy, we investigate character mixing— the unintended production of Simplified Chinese characters under Taiwanese-Mandarin-style prompts—and propose a human-grounded purity penalty derived from perceptual thresholds measured from users. Our evaluation reveals substantial character contamination (3%–19%) across state-of-the-art VLMs. We find that Gemini-3-Pro significantly outperforms the strongest open-weight baseline, Qwen3 235B MoE, by up to 22 percentage points on dialogue tasks once the purity penalty is applied. These results highlight orthographic consistency as a vital, yet overlooked, dimension for localized multimodal evaluation and deployment.