← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Eye-based Emotion Recognition via Event-Driven Sparse Transformers

Zixuan Wan, Jiqing Zhang, Yushan Wang, Hu Lin, Yafei Wang 0004, Zetian Mi, Xin Yang 0011, Xianping Fu, Huibing Wang

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755551 ↗

摘要

Event-driven eye-based emotion recognition has attracted increasing attention due to the high temporal resolution and dynamic range inherent to event cameras. The intrinsic spatial sparsity of event data, combined with the eye-based emotion recognition task's reliance on localized features such as eyebrows and eyelids, makes it intuitive and efficient to discard less informative regions. However, integrating such sparsification into CNNs remains challenging due to their reliance on dense grid-based operations. In this paper, we propose an efficient vision transformer framework for eye-based emotion recognition with event cameras. Specifically, we present window selection and token selection schemes tailored for event data and eye-based emotion recognition, which can diminish computing demands while enhancing performance. Firstly, we estimate the importance of all local windows and discard those with limited information, reducing computational cost while emphasizing attention on the periocular region. Secondly, we further introduce an adaptive token pruning mechanism that jointly evaluates the input event data and tokens to predict a binary decision mask, identifying and discarding uninformative tokens. Extensive experiments validate that the proposed approach outperforms existing state-of-the-art methods in accuracy by a significant margin.