Large-Small Model Synergy with Multimodal Fine-Grained Heuristics for Knowledge-Based Visual Question Answering
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754750 ↗
摘要
Multimodal Large Language Models (MLLMs) possess extensive knowledge and strong reasoning capabilities, achieving remarkable performance in knowledge-based visual question answering, significantly surpassing traditional small-scale Vision-Language Models (VLMs). However, the distinct training paradigms of MLLMs and small-scale VLMs result in misaligned feature representation spaces and divergent answer prediction distributions. To bridge this gap, we propose a novel end-to-end large-small model synergy framework, where small VLMs and MLLMs collaborate via synergistic optimization of shared objectives while maintaining their co-evolving complementary specializations. Specifically, multimodal fine-grained heuristics are extracted from well-tuned small VLMs and subsequently projected into the textual space of MLLMs through dedicated visual and textual collaboration modules. This enables cross-modal guidance for both visual and textual inputs. Finally, a dual-objective synergy loss promotes alignment toward shared goals, while a visual discrepancy loss preserves specialization diversity. Extensive experiments demonstrate that our framework achieves state-of-the-art performance on both the OK-VQA and A-OKVQA benchmarks.