← 返回论文检索
ACM Multimedia 2025Content: Multimodal Fusion

DeepMolTex: Deep Alignment of Molecular Graphs with Large Language Models via Mixture of Modality Experts

Mingliang Yan, Yanhua Yu, Ruochi Zhang, Zhiyuan Liu 0010, Ruicheng Zhang, Yimeng Ren 0001, Kangkang Lu 0002, Zhiyong Huang 0010, Feng Luo 0004, Zhen Cai

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755875 ↗

摘要

Recent advances in Molecular Graph-Language Models (MGLMs) have demonstrated promising capabilities in molecular understanding tasks. However, existing approaches face critical limitations: (1) shallow alignment methods which employ identical processing modules for both modalities, resulting in compromised expressiveness and catastrophic forgetting of pre-trained language capabilities; and (2) over-reliance on high-level molecular representations that inadequately capture fine-grained structural information essential for comprehensive molecular understanding. To address these challenges, we present DeepMolTex, a novel framework for Deep fusion of Molecular structure and Textual representations across multiple scales. Our approach introduces a Mixture of Modality Experts (MoME) architecture that facilitates deep alignment between molecular graph features and large language models while preserving language capabilities, and a multi-scale graph projector that extracts and aligns molecular features at atom, motif, and molecule levels. Experimental results demonstrate that DeepMolTex significantly outperforms existing methods on fundamental molecular understanding tasks, including molecule description generation and IUPAC name prediction, while effectively preserving the language capabilities of the pre-trained LLM.