Fast-SAM3D: 3Dfy Anything in Images but Faster
Institute of Computing Technology, Chinese Academy of Sciences; University of Chinese Academy of Sciences · University of the Chinese Academy of Sciences · China University of Mining Technology - Beijing · Institute of Computing Technology, Chinese Academy of Sciences · ETH Zurich · CUNY City College of NY · Shanxi University · Shanghai Jiao Tong University · ETHZ - ETH Zurich
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
SAM3D enables scalable, open-world 3D reconstruction from complex scenes, yet its deployment is hindered by prohibitive inference latency. In this work, we conduct the **first systematic investigation** into its inference dynamics, revealing that generic acceleration strategies are brittle in this context. We demonstrate that these failures stem from neglecting the pipeline's inherent multi-level **heterogeneity**: the kinematic distinctiveness between shape and layout, the intrinsic sparsity of texture refinement, and the spectral variance across geometries. To address this, we present **Fast-SAM3D**, a training-free framework that dynamically aligns computation with instantaneous generation complexity. Our approach integrates three heterogeneity-aware mechanisms: (1) *Modality-Aware Step Caching* to decouple structural evolution from sensitive layout updates; (2) *Joint Spatiotemporal Token Carving* to concentrate refinement on high-entropy regions; and (3) *Spectral-Aware Token Aggregation* to adapt decoding resolution. Extensive experiments demonstrate that Fast-SAM3D delivers up to **2.67$\times$** end-to-end speedup with negligible fidelity loss, establishing a new Pareto frontier for efficient single-view 3D generation.