← 返回论文检索
ACM Multimedia 2025Content: Vision and Language

Visual Localization using Hybrid Feature Grid and Learned Weighted Global Point Cloud

Junyi Wang 0001, Yue Qi

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755361 ↗

摘要

To fully leverage diverse scene representations for visual relocalization, we propose a novel localization framework that systematically establishes inter-frame relationships and integrates multiple feature modalities. Our localization pipeline comprises three key stages, containing initial pose estimation using local point cloud structure, pose refinement by hand-crafted features and 3D Gaussians, and pose confidence estimation through a leaned global representation. Specifically, the initial stage begins with aligning a known source point cloud to a predicted local Target Point Cloud (TPC) using a registration algorithm. For pose refinement, we introduce the Hybrid Feature Grid (HFG), which fuses hand-crafted points and 3D Gaussians to enrich texture cues. To assess pose reliability, we propose the learned Weighted Global Point Cloud (WGPC), aggregating multi-frame information to enhance confidence estimation. To jointly learn TPC, HFG, and WGPC, we design a Siamese Localization Network (SiaLocNet) featuring three core innovations, including learning trajectory-based features for the limitation of single-view inputs, a feature fusion module to facilitate the construction of the three core structures. and an inverse self Chamfer Distance along with a shape-aware term to improve the robustness of WGPC. Extensive experiments on the 7 Scenes and Cambridge Landmarks datasets demonstrate that our method achieves state-ofthe-art performance across both indoor and outdoor environments.