← 返回论文检索
ACM Multimedia 2025Experience: Multimedia Applications

Bridging Inter-Class Ambiguity and Spatial Variability in Flexible Object Recognition via Graph Distillation

Lin Zuo, Kunshan Yang, Mengmeng Jing, Xiangxu Zhao, Jiaqiao Chen

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3755858 ↗

摘要

Flexible object recognition remains challenging in multimedia scenarios due to inherently diverse shapes and sizes, and subtle inter-class differences. Graph-based vision models show promise in flexible objects recognition by capturing variable relationships. However, they suffer from two problems: (1) inter-class ambiguity hinders model discrimination and (2) frequent scale changes degrade model generalization. To address these limitations, we propose a unified graph distillation framework that enhances inter-class discrimination and spatial generalization while maintaining computational efficiency. For inter-class ambiguity problem, we introduce a virtual prototype module that dynamically generates learnable class prototypes via clustering intermediate features. These prototypes are incorporated into the distillation loss to sharpen decision boundaries. A global-local distillation mechanism further capture both image-level global semantics and patch-level local details, enhancing inter-class discrimination. For frequent scale changes problem, we design a patch-aware distillation strategy that transfers knowledge across multiple patch scales, strengthening the student model's spatial generalization to match various shapes and sizes of flexible objects, thus alleviate generalization degradation. Extensive experiments on flexible-object datasets (FDA, FSCW, CCSN) and challenging benchmarks (CIFAR-100, Mini-ImageNet) confirm effectiveness and efficiency of our method.