TongGu-VL: Advancing Visual-Language Understanding in Chinese Classical Studies through Parameter Sensitivity-Guided Instruction Tuning
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755769 ↗
摘要
Chinese Classical Studies (CCS) is a pivotal gateway to ancient Chinese culture. Spanning ancient texts, illustrations, paintings, and calligraphy, CCS presents significant challenges for non-specialists due to its language and visual complexity. While Large Language Models (LLMs) have been explored to facilitate CCS, current methods primarily focus on textual analysis, overlooking the rich visual information intrinsic to classical materials. To bridge this gap, we propose TongGu-VL, a pioneering specialized MLLM designed for CCS applications. Our contributions are threefold. First, we construct CCS358K, a comprehensive multimodal instruction dataset to enhance MLLMs' CCS capabilities. Second, we propose Parameter Sensitivity-Guided Instruction Tuning (PSG-IT), a novel method that mitigates catastrophic forgetting without data replay. It effectively preserves TongGu-VL's general skills, while optimizing its CCS performance. Third, we design a Visual-Text Early Fusion (VTEF) module, which harnesses MLLMs' modality alignment to generate instruction-aware visual representations, thus improving language modeling. Extensive experimental results demonstrate that our model outperforms existing MLLMs on a broad range of CCS tasks, while maintaining general capabilities that benefit other domains beyond CCS. Our model and dataset will be publicly available.