← 返回论文检索
CVPR 2026

SCIEval: Evaluating and Benchmarking the Faithfulness of Scientific Image Generation and Interpretation with Large Multimodal Models

Guanghui Ye, Huan Zhao, Zhixue Zhao, Tengfei Ma, Kehan Wang, Steffen Eger, Zhihua Jiang

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Scientific images often require accurate numerical representations and correct object attributes. However, current faithfulness metrics are primarily tailored toward photorealistic, real-life imagery, rendering them ill-suited for scientific image evaluation. To address this gap, we introduce a novel evaluation model, SCIEval (Scientific Image Evaluation), which aims to capture faithfulness through three key dimensions: (i) Relevance, measuring overall text-image correspondence; (ii) Accuracy, examining the technical details of scientific objects; and (iii) Explainability, which isolates unfaithful elements within the generated content. To address these dimensions, we curate a specialized dataset of scientific text-image pairs to train three evaluation modules. For the Relevance and Accuracy modules, we propose a CLIP-based strategy that enhances scientific image perception through intra- and cross-modal contrastive learning. Concurrently, the Explainability module is developed by fine-tuning a high-performance Large Multimodal Model (LMM) using supervised rationale signals. Finally, we present SCIEval-Bench, a human-annotated evaluation benchmark consisting of 3,000 samples for scientific text-to-image and 3,000 samples for scientific image captioning. Extensive experiments on SCIEval-Bench demonstrate that our SCIEvalmodel is significantly more reliable than 24 competing models--including GPT-4o--exhibiting a superior correlation with human judgments.