OmniDoctor: Towards LLM-centric Lifelong Learning for New Emerging Medical VQA Tasks
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755745 ↗
摘要
Current medical multimodal large language models (MLLMs) have demonstrated high accuracy and effectiveness on specific medical visual question answering (Medical VQA) tasks. However, they largely fail to tackle continuously emerging unseen Medical VQA scenarios (e.g., MRI, X-ray) in real-world settings, which significantly hinders their broader adoption in practical clinical environments. Motivated by these gaps, this paper introduces a new task, namely LLM-centric Lifelong Learning for New Medical VQA (L3NMV), which enables Large Language Models (LLMs) to continually learn medical image-text knowledge across various medical VQA tasks. Furthermore, this paper reveals two critical challenges: 1) Efficient medical knowledge retention (Each-task), which aims to retain essential knowledge for each Medical VQA task efficiently with limited data. 2) Efficient medical interference mitigation (Cross-task), which focuses on efficiently mitigating information interference across various Medical VQA tasks with knowledge barriers. To address these challenges, this paper proposes the OmniDoctor model, i.e., an omniscient doctor that simulates how doctors continuously update their knowledge and skills through continuous medical education, with the goal of equipping the model with lifelong learning capabilities via an efficient incremental medical parameter constraining mechanism for L3 NMV. This model is designed with two key modules to address the above two challenges, respectively. Especially, this paper constructs an Unseen L3NMV dataset to simulate real-world incremental clinical scenarios. Extensive experiments on this dataset demonstrate that OmniDoctor outperforms several advanced lifelong learning baselines. These results justify the significance of the L3 NMV task and the effectiveness of OmniDoctor in continually adapting to new Medical VQA tasks.