← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

OCR-Critic: Aligning Multimodal Large Language Models' Perception through Critical Feedback

Qiuna Tan, Runqi Qiao, Guanting Dong 0001, Yifan Zhang, Minhui Wu, Jiapeng Wang 0005, Miaoxuan Zhang, Yida Xu, Chong Sun, Chen Li 0031, Honggang Zhang 0002

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3754585 ↗

摘要

Recent advancements in Large Multimodal Models have demonstrated impressive performance in various tasks. However, their capabilities in error detection and resolution for Optical Character Recognition (OCR) remain underexplored. To address this gap, we construct the first visual instruction tuning dataset specifically for detailed OCR error analysis. Building on this foundation, we develop a universal, plug-and-play OCR-Critic model that incorporates three novel dynamic alignment strategies. These strategies systematically mitigate LMMs' weaknesses in OCR tasks by providing coarse-to-fine error feedback. To comprehensively evaluate these capabilities, we introduce OCR-ERROR, a benchmark designed to assess LMMs' ability to detect and categorize OCR errors, covering two task types, diverse error categories, and 2,400 rigorously validated samples. Experimental results show that OCR-Critic effectively identifies fine-grained OCR errors across multiple domains. With the integration of our dynamic alignment strategies, the LMM further achieves substantial performance gains on four prominent benchmarks, demonstrating both versatility and effectiveness.