← 返回论文检索
ACM Multimedia 2025Content: Media Interpretation

Retaining Temporal Semantics and Relation Topologies for Continual Weakly-Supervised Audio-Visual Video Parsing

Jie Fu 0004, Bingkun Bao

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755561 ↗

摘要

To achieve audio and visual action detections in a given video accompanying with only video-level labels, a group of weakly-supervised audio-visual video parsing methods have been explored. Throughout their training processes, the action categories are typically assumed to be static, which is not always satisfied. Consequently, these methods can not be employed to handle dynamic scenarios involving continuously growing novel classes. To alleviate the above issue, we introduce a novel Continual Weakly-Supervised Audio-Visual Video Parsing (C-WSAVVP) task, where maintaining the knowledge of historic categories remains the eternal topic. Distinctly, owing to the weakly-supervised and multi-modal characteristics, two core challenges are more obvious in C-WSAVVP: (1) Compared with the continual audio-visual video classification task, where distilling video-level coarse action semantic of trimmed videos is sufficient for mitigating catastrophic forgetting, C-WSAVVP has to retain more fine-grained temporal semantic information of untrimmed videos containing both actions and backgrounds. (2) The semantics of different actions generally exhibit a certain degree of correlation, which is beneficial for understanding related actions, but how to maintain the semantic correlations? To address the specific challenges, the Semantic Prototype-based Action Refinement (SPAR) and Inter-Class Relation Topology Preservation (IRTP) modules are explored, where the former devotes to utilizing various semantic prototypes to refine more reliable temporal action intervals for distillation and the latter focuses on retaining semantic correlations between different actions in both modalities during continual learning. Comprehensive experiments on our reconstructed C-LLP dataset demonstrate the effectiveness and generalization capability of our proposed method.