← 返回论文检索
ICLR 2026PosterAccept (Poster)

FlexRibbon: Joint Sequence and Structure Pretraining for Protein Modeling

Jianwei Zhu, Yu Shi, Ran Bi, Peiran Jin, Chang Liu, Zhe Zhang, Haitao Huang, Zekun Guo, Pipi Hu, Fusong Ju, Lin Huang, Xinwei Tai, Chenao Li, Kaiyuan Gao, Xinran Wei, Huanhuan Xia, Jia Zhang, Yaosen Min, Zun Wang, Yusong Wang, Liang He, Haiguang Liu, Tao Qin

Zhongguancun Academy · Beijing Zhongguancun Academy · Microsoft Research · Zhongguancun Institute of Artificial Intelligence · Hunan University · Tsinghua University, Tsinghua University · Beijing University of Post and Telecommunications · Huazhong University of Science and Technology · The Hong Kong University of Science and Technology(Guang Zhou) · Microsoft Research AI for Science · Microsoft · Shanghai Artificial Intelligence Laboratory · Xi‘an Jiaotong University · Zhongguancun Acadamy · Microsoft Research · Microsoft Research Asia

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Protein foundation models have advanced rapidly, with most approaches falling into two dominant paradigms. Sequence-based language models (e.g., ESM-2) capture sequence semantics at scale, and a number of recent works incorporate structural signals into sequence encoders. MSA-based predictors (e.g., AlphaFold 2/3) achieve accurate folding by exploiting evolutionary couplings, but their reliance on homologous sequences makes them less reliable in highly mutated or alignment-sparse regimes. We present FlexRibbon, a pretrained protein model that jointly learns from amino acid sequences and three-dimensional structures. Our pretraining strategy combines masked language modeling with diffusion-based denoising, enabling bidirectional sequence-structure learning without requiring MSAs. Trained on both experimentally resolved structures and AlphaFold 2 predictions, FlexRibbon captures global folds as well as flexible conformations critical for biological function. Evaluated across diverse tasks spanning interface design, intermolecular interaction prediction, and protein function prediction, FlexRibbon establishes new state-of-the-art performance on 12 different tasks, with particularly strong gains in mutation-rich settings where MSA-based methods often struggle.