Foundation Encoders Are All You Need for Preference-Aware Personalization
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Personalized image generation based on user behavior reflects individual preferences with minimal user intervention. However, existing studies often rely on inaccurate profiling, high resource costs, and model-specific designs, which jointly restrict creativity, diversity, and generality. To address these limitations, we propose FAN, a novel approach that enables preference-aware personalization using only foundation encoders, without additional structures. FAN performs tailored profiling to capture user preferences and reconstructs transformer-based encoders to integrate them while preserving target fidelity. Experiments show that FAN achieves robust personalization across various foundation text-to-image models. It also extends to applications such as CLIP retrieval, unCLIP, and vision-language models, seamlessly integrating into diverse encoders without fine-tuning.