← 返回论文检索
ACM Multimedia 2024Oral Session 26: Cultural Heritage & Media Analysis

Cognition-Supervised Saliency Detection: Contrasting EEG Signals and Visual Stimuli

Jun Ma, Tuukka Ruotsalo

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3664647.3681037 ↗

摘要

Understanding human assessment of semantically salient parts of multimedia content is crucial for developing human-centric applications, such as annotation tools, search and recommender systems, and systems able to generate new media matching human interests. However, the challenge of acquiring suitable supervision signals to detect semantic saliency without extensive manual annotation remains significant. Here, we explore a novel method that utilizes signals measured directly from human cognition via electroencephalogram (EEG) in response to natural visual perception. These signals are used for supervising representation learning to capture semantic saliency. Through a contrastive learning framework, our method aligns EEG data with visual stimuli, capturing human cognitive responses without the need for any manual annotation. Our approach demonstrates that the learned representations closely align with human-centric notions of visual saliency and achieve competitive performance in several downstream tasks. We also introduce an open EEG/image dataset to facilitate research in utilizing cognitive signals for multimodal data analysis and developing models for cross-modal representation learning.