MLLMs Meet Person Re-identification
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3758150 ↗
摘要
Person re-identification (Re-ID) models have achieved remarkable advancements with the advent of deep learning. However, their performance often degrades in diverse scenarios, such as variations in viewing angles, lighting conditions, and environmental changes. These limitations arise from the difficulty in generalizing across multiple factors, including environments and subject appearances. Multimodal Large Language Models (MLLMs) offer a promising alternative to address these challenges by leveraging generalized knowledge, as demonstrated in biometric tasks like face and iris recognition. This study explores the Re-ID capabilities of MLLMs by comparatively evaluating six representative MLLMs on the most challenging scenarios, including angle variation, illumination differences, clothing changes, image corruption, and visually fine-grained scenarios in Re-ID. We find that GPT-4o outperforms other MLLMs in handling angle variation, illumination differences, corruption resistance, and fine-grained detail disturbances, demonstrating high accuracy and robustness in challenging Re-ID scenarios. However, further optimization is required for robustness against illumination variation, corruption handling, and fine-grained identification across all tested MLLMs. Additionally, the Re-ID performance of MLLMs can be improved by applying several prompt templates. Our research suggests potential directions for integrating MLLMs into Re-ID systems to enhance performance and robustness, underscoring their promising potential in this field.