DSP: Dense-Sparse Parallel Networks for Self-supervised 3D Multi-person Pose Estimation from Multiple Views
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755531 ↗
摘要
Recent advances in multi-view 3D multi-person pose estimation have led to significant progress. However, several critical challenges remain, including the limited extraction and integration of multi-domain information, as well as the high annotation costs associated with 3D data in multi-person scenarios. These issues hinder the broader applicability of current methods in complex computer vision tasks. In this paper, we propose a Dense-Sparse Parallel Networks (DSP) framework that jointly leverages spatial, temporal, and frequency-domain information through an adaptive geo-consistency self-supervised strategy. Specifically, we design a multi-view spatial feature extraction module that captures cross-view spatial distributions from dense multi-view feature maps. In parallel, we employ a local-global temporal attention module and a frequency-aware attention module to extract dynamic temporal patterns and localized frequency-domain features from sparse keypoint data. Furthermore, a multi-domain parallel fusion module is introduced to effectively integrate features across all domains, enabling accurate multi-person 3D pose regression. To enhance self-supervised learning, we employ a dynamic view selector guided by reinforcement learning, which reduces the impact of inaccurate pre-trained 2D poses. Experimental results on three benchmark datasets (i.e., CMU Panoptic, Campus, and Shelf) demonstrate that the proposed DSP framework achieves robust and accurate performance, as evidenced by comparisons with other state-of-the-art methods.