← 返回论文检索
NeurIPS 2025{location} PosterAccept (poster)

How Many Tokens Do 3D Point Cloud Transformer Architectures Really Need?

Tuan Tran Anh, Duy M. H. Nguyen, Hoai-Chau Tran, Michael Barz, Khoa D Doan, Roger Wattenhofer, Vien Ngo, Mathias Niepert, Daniel Sonntag, Paul Swoboda

German Research Center for AI · DFKI & Max Planck Research School for Intelligent Systems · University of Illinois at Urbana-Champaign · German Research Center for Artificial Intelligence, DFKI · VinUniversity · ETH Zurich · Bosch Center for Artificial Intelligence · Universität Stuttgart and NEC Labs Europe · Heinrich-Heine University Düsseldorf

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Recent advances in 3D point cloud transformers have led to state-of-the-art results in tasks such as semantic segmentation and reconstruction. However, these models typically rely on dense token representations, incurring high computational and memory costs during training and inference. In this work, we present the finding that tokens are remarkably redundant, leading to substantial inefficiency. We introduce an efficient token merging method and illustrate that it can reduce the token count by up to 90–95% while maintaining competitive performance. This finding challenges the prevailing assumption that more tokens inherently yield better performance and highlights that many current models are over-tokenized and under-optimized for scalability. We validate our method across multiple 3D vision tasks and show consistent improvements in computational efficiency. This work is the first to assess redundancy in large-scale 3D transformer models, providing insights into the development of more efficient 3D foundation architectures. Our code and checkpoints are publicly available at https://gitmerge3d.github.io.