AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
We present AMB3R, a multi-view feed-forward model for metric-scale dense 3D reconstruction that addresses diverse 3D vision tasks. The key idea is to employ a sparse, yet compact, volumetric scene representation as our backend, enabling geometric reasoning with spatial compactness. We further introduce AMB3R-VO and AMB3R-SfM, two training-free, model-agnostic pipelines that extend multi-view networks to uncalibrated visual odometry (online) or structure from motion for any number of images without test-time optimization. Compared to prior 3D foundation models, AMB3R achieves state-of-the-art performance in camera pose, depth, metric-scale estimation, and 3D reconstruction, while AMB3R-VO and AMB3R-SfM even surpass optimization-based SLAM and SfM systems with dense reconstruction priors on common benchmarks.