Ground and Reconstruct: Entity-Region Bidirectional Alignment Pre-Training for Low-Resource GMNER
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755424 ↗
摘要
Grounded Multimodal Named Entity Recognition (GMNER) extends Multimodal Named Entity Recognition (MNER) by identifying named entities, their types, and corresponding image regions. Fine-grained MNER and Grounding (FMNERG) further refines entity categorization. However, existing methods struggle with scarce annotated data, particularly in low-resource scenarios, and often fail to generalize to unseen entities. While vision-language pre-training (VLP) leverages unlabeled image-caption pairs, it primarily learns generic visual-linguistic representations, overlooking fine-grained entity-region alignment crucial for entity-related tasks. To address these challenges, we propose a unified VLP framework for GMNER and FMNERG, introducing two task-specific pre-training objectives: Entity-to-Region Alignment (ETRA) for entity grounding and Region-to-Entity Alignment (RTEA) for entity reconstruction. These tasks jointly optimize fine-grained entity-region alignment. To compensate for the lack of fine-grained multimodal pre-training data, we develop an automatic labeling method that distills entity-oriented knowledge from large-scale unlabeled image-text pairs, enhancing generalization to unseen entities. Extensive experiments on GMNER and FMNERG benchmarks demonstrate that our framework outperforms existing low-resource learning approaches and achieves competitive performance in full-supervision, underscoring its effectiveness across diverse data conditions.