← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Visual Perception Uncertainty Learning for Hallucination Detection in Large Vision-Language Models

Runze Zhao, Fuqing Zhu, Jizhong Han, Songlin Hu 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755126 ↗

摘要

Hallucination remains a significant challenge which constrains the development of large vision-language models (LVLMs). Therefore, reliable hallucination detection has become a critical step in LVLMs evaluation and real-world deployment. Many previous studies have explored hallucination detection in LVLMs, with uncertainty-based approaches being widely adopted due to the independence from external tools and relatively low resource consumption. However, we observe that uncertainty does not always completely correlate with hallucination. Therefore, uncertainty-based methods may fail in certain cases, such as instances exhibiting high uncertainty but non-hallucination. To address this issue, we propose a framework called Visual Perception Uncertainty Learning (VisPUL) for hallucination detection in LVLMs. Specifically, VisPUL integrates visual information into uncertainty learning directly, allowing to capture uncertainty and visual-text consistency simultaneously. VisPUL improves the insufficiency of uncertainty methods that rely only on text output, providing enhanced generalizability and reliability. Extensive experiments conducted on the M-HalDetect and POPE datasets, covering both open-ended and yes-or-no tasks. Experimental results demonstrate that VisPUL significantly outperforms several strong baseline methods across different LVLMs.