← 返回论文检索
ACM Multimedia 2024Grand Challenges

Enhancing Multimodal Large Language Models on Demonstrative Multi-Image Instructions

Xian Fu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3664647.3688994 ↗

摘要

Recent advancements in Multimodal Large Language Models (MLLMs) have showcased their ability to handle tasks involving single images, such as generating detailed descriptions and answering related queries. These models have significantly pushed the boundaries of multimodal content comprehension and generation. However, when confronted with more complex multimodal contexts, particularly those involving interleaved text and multiple images, MLLMs often face difficulties in effectively processing and understanding these complex contexts. Furthermore, MLLMs exhibit weaknesses in instruction following, especially when dealing with tasks that require reasoning about and prioritizing specific details. In this paper, we design a novel approach, InstructFusion, which introduces an enhanced module Fusing Former to improve the model's understanding of relationships among multiple images within convoluted multimodal tasks. This ensures more accurate and contextually relevant responses. Additionally, we adopt a two-stage training process, first developing general capabilities in the Fusing Former and then fine-tuning with LoRA on instruction-focused datasets to enhance instruction following ability while minimizing costs and preventing forgetting. Empirical results, including a second-place finish in the ACM Multimedia 2024 Demonstrative Instruction Following Challenge, demonstrate the effectiveness of our proposed method.