Referring Multi-Object Tracking in Satellite Videos: A New Benchmark and Baseline
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3758254 ↗
摘要
Referring multi-object tracking (RMOT), which aims to track one or more objects in a video based on a natural language query, is increasingly crucial for a wide range of real-world applications. However, the study of RMOT in satellite video (RMOT-SV) scenarios remains limited, largely due to the high cost of data acquisition and the difficulty of annotation. To address this gap, we introduce RefSat, the first dataset for benchmarking RMOT-SV. RefSat comprises 212 top-down viewpoint video clips, totaling 31,129 frames, collected from a variety of publicly available satellite video datasets. By combining manual annotations of object appearance and position with automatic motion estimation, we build a semi-automatic pipeline that generates high-quality natural language descriptions covering object attributes and motion trajectories, resulting in over 4,000 objects paired with carefully designed textual queries. RefSat features satellite-specific challenges such as small object sizes, cloud occlusions, and motion-referenced semantics. To address these, we introduce RSRefTrack, a tailored baseline designed for small object perception and motion-aware grounding, which outperforms existing state-of-the-art RMOT methods on the RefSat benchmark. Project page: https://github.com/Zhang-Peirong/RefSat