← 返回论文检索
CVPR 2026

STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection

Huijie Fan, Pengrui Huang, Qiang Wang, Baojie Fan, Jiahua Dong, Liangqiong Qu

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Existing surrounding-view 3D object detectors initialize high-confidence queries using current 2D information, while leveraging historical 3D features as priors. However, such heavy reliance on 2D cues introduces spatio-temporal inconsistencies between 2D and 3D representations. Specifically, 2D cues lack sufficient spatial information, limiting 3D localization capability. Moreover, insufficient temporal interaction often leads to object omission under occlusion. To address these challenges, we propose STUR3D, a unified framework establishing spatio-temporal alignment between 2D and 3D perception. First, we project temporal 3D features to the 2D image plane, empowering the 2D detector to distill representations essential for 3D localization, harmonizing cross-dimensional information. Second, we inject temporal cues into 2D detection, fostering spatio-temporal reasoning, and ensuring robust 3D detection under dynamic scenes and occlusion. Additionally, we embed depth-aware geometric cues into features for 2D-to-3D lifting, mitigating inherent ambiguities. Extensive nuScenes experiments validate STUR3D, achieving SOTA on the test set with 57.9% mAP and 64.6% NDS.