Video-based Transparent Object Segmentation via Temporal Feature Aggregation
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755084 ↗
摘要
Transparent object segmentation from a single image has been investigated for several years. However, detecting transparent areas from video has not been well explored, especially for different kinds of transparent categories besides glass, due to the scarcity of such a dataset. Therefore, in this paper, we propose the video-based transparent object segmentation task and introduce the first-of-its-kind corresponding dataset named TransVid, which contains nearly 400 videos with a total of 18,523 frames. Based on TranVid, we further propose a new method called TranSeg, in which we innovatively introduce Graph Neural Networks into the temporal segmentation task and combined with a novel Diffusion Model to make the model's segmentation results more accurate. Experimental results show that TranSeg achieves higher accuracy with fewer parameters than previous state-of-the-art models, demonstrating the effectiveness of our method. Moreover, comprehensive ablation analysis reveal several fascinating insights and suggest viable paths for further research.