← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Formula Spotting Based on Synergy Perception and Representation Mining

Gang Pan 0002, Hongen Liu, Di Sun 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755773 ↗

摘要

Formula spotting aims to simultaneously detect and recognize formulas in documents, with broad applications in intelligent document parsing, mathematical reasoning, and more. Although existing methods that first detect and then recognize have achieved prominent results, they still suffer from semantic confusion from similar character structures, semantic loss from bounding box perturbation, and visual interference from non-formula regions. To address these issues, we propose a Synergy Perception and Representation Mining Network. This network facilitates explicit interaction between the RoI features of the detection module and the semantic features of the recognition module, leveraging additional visual priors to distinguish subtle differences in similar characters. Moreover, to better perceive the boundary character structure of formulas and filter out irrelevant visual interference, a Formula Representation Mining module is proposed. This module employs progressive attention mining to achieve a complementarity between semantic information and visual context without disrupting the linguistic priors of the formulas. Additionally, to enhance the efficiency of formula decoding, we propose a parallel mask, allowing the network to output multiple LaTeX tokens simultaneously in a single prediction step. To evaluate the effectiveness of our method in formula spotting, two novel datasets: Formula-7K and Exam-1K are established. To the best of our knowledge, they are the first formula spotting datasets. Experimental results on Formula-7K and Exam-1K validate the generality and effectiveness of the proposed method. Code is available at https://github.com/hongen123/SynRMFormer.