← 返回论文检索
ACM Multimedia 2025Content: Multimodal Fusion

Rule Meets Learning: Confidence-Aware Multi-View Fusion for Self-Supervised 3D Hand Pose Estimation

Pengfei Ren 0001, Jingyu Wang 0001, Haifeng Sun 0001, Qi Qi 0001, Jing Wang 0039, Jianxin Liao

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755342 ↗

摘要

Self-supervised 3D hand pose estimation methods can leverage labeled synthetic data along with unlabeled real-world data for model training, thereby alleviating the reliance on large-scale annotated datasets. Multi-view information fusion is a key factor in the success of these methods. Rule-based fixed fusion methods are simple, efficient, and generalizable, but they neglect the rich visual information in each view. Neural network-based learnable fusion methods can effectively model both intra- and inter-view semantic context, but they tend to overfit to the domain-specific feature of synthetic data and susceptible to interference of domain gaps. In this paper, we decompose multi-view fusion into two components: a learnable confidence estimation stage and a fixed confidence fusion stage. This design not only enables effective use of multi-view semantic cues but also ensures strong cross-domain generalization. To achieve accurate and robust confidence estimation, our method jointly exploits both multi-view pose consistency and pose-to-data consistency. Experiments on three public datasets demonstrate that our approach significantly outperforms existing state-of-the-art self-supervised 3D hand pose estimation methods.