Cross-Counter-Repeat Attention for Enhanced Understanding of Visual Semantics in Radiology Report Generation
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755368 ↗
摘要
Radiology report generation (RRG), intended to automatically generate a coherent free-text report describing the clinical observations of a radiograph, has been attracting increasing attention from researchers. In recent years, the Transformer-based encoder-decoder architecture has been adopted by most existing methods. However, they neglect the structural rationality issue when applying this single-modal architecture to the multi-modal RRG task, where information can only flow from visual features to textual features, but not in the opposite direction. This information asymmetry results in visual features having no knowledge of the textual features, sending out all visual information, including a large amount of heterogeneous noise. Consequently, this introduces significant resistance to the downstream decoder, which substantially limits or even harms the generation process. To tackle this problem, we present a method where a cross-counter-repeat attention is developed to integrate useful information from two separate modalities, and a memory-driven visual semantics enhancing module is designed to reinforce the visual features with strong time-ordered semantic information. Experimental results on the widely-used IU-Xray dataset show that our approach achieves the state-of-the-art performance, with a remarkable 6.9% improvement in BLEU-4 score. Further analyses also demonstrate that our method can generate sufficiently comprehensive reports to assist radiologists in their clinical decision-making.