End-to-End Multiple Object Tracking with Dynamic Scene Perception
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755617 ↗
摘要
End-to-end Multiple Object Tracking (MOT) frameworks integrate detection and tracking into a unified model, avoiding intermediate information loss and complicated post-processing. However, existing end-to-end MOT trackers rely on track queries of the previous frame to provide prior information. Their limited short-term temporal modeling struggle to cope with high dynamic tracking scenarios, where inter-frame target variations exhibit significant heterogeneity. To address these shortcomings, we propose a scene-perception MOT framework (SP-MOT) that encodes scene context understanding into long-term embedding and adaptively complements it with short-term cues, enabling discriminative and flexible instance representations. Specifically, SP-MOT introduces: (1) a learnable scene query that globally profiles foreground and background to capture short-term scene-level features; (2) a context understanding module to uncover long-term stable relationships across dynamic scenes based on multiple historical scene features; (3) scene-adaptive augmented decoding that leverages scene information as guidance, adaptively aggregating long-term and short-term information into object embeddings, improving the model's association ability and fault-tolerance. Extensive experiments on MOT benchmarks demonstrate that SP-MOT outperforms state-of-the-art end-to-end trackers across multiple metrics, particularly in challenging scenarios with high dynamics.