← 返回论文检索
ACM Multimedia 2025Content: Media Interpretation

Cause and Effect: Video Social Relationship Recognition from Causal Perspective

Yuxuan Zhang, Bo Wang, Yu Du, Yangfu Zhu, Haorui Wang, Guangyao Su, Tao Zhou, Bin Wu 0001

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755732 ↗

摘要

Video social relation recognition is a fundamental task in video understanding, which is dedicated to the construction of multi-modal knowledge graphs. Previous work mainly focuses on multi-modal fusion and the construction of special character graphs. However, they often treat the global frame sequence equally, ignoring the influence of key frame sequence on relation recognition. Specifically, the key frame sequence that significantly reflect character relationships in a video tends to be sparse and short. At the same time, the key frames have not only temporal but also strong causal relationship. Therefore, we propose a novel Video Local Causal Frame (VLCF) model to explore the causal relationship between frames. Inspired by Granger causality theory, we estimate inter-frame causal relationships by comparing the predicted result frames with and without masking the premise frame. We then construct global connections between video frames. Multiple local causal frame sequences and global frame sequences are extracted to capture the key information and global information in the video. Extensive experiments conducted on the ViSR dataset and the MovieGraphs dataset demonstrate that the proposed model achieves state-of-the-art performance.