← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image Captioning

Yuxin Xie, Dongyue Chen 0001, Yue Zhu, Tong Jia 0001, Shizhuo Deng

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755762 ↗

摘要

Image captioning aims to create natural language descriptions of images. Recent advancements in image captioning have explored text-only training methods that eliminate the need for image annotations. However, these methods are prone to generate descriptions that include objects that do not actually appear in the image, but are instead drawn from the retrieval texts or hard prompts-resulting in object hallucinations. To address this issue, we propose synergistic prompting mechanism called NASCap (Noise-Aware Decoding with Salient Region Enhancing for Zero-Shot Image Captioning). Our method further improves the accuracy of generated captions by designing a fusion model with mask attention that isolate integrates retrieved captions with input features. The hard prompt mixed with negative entities is designed to further improve model's robust to wrong information. Additionally, we introduce a training-free multi-granularity fusion strategy that dynamically perceive and enhance salient regions into global representation. Extensive experiments demonstrate that NASCap sets a new state-of-the art cross-domain (transferable) captioning and performs Through extensive experiments, our straightforward yet powerful approach has demonstrated its efficacy, outperforming the state-of-the-art methods by a significant margin in image captioning compared to zero-shot captioning based on text-only training.