← 返回论文检索
ICLR 2026PosterAccept (Poster)

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

Chengzhi Mao, Xudong Lin, Wen-Sheng Chu

Rutgers University · Columbia University · Google Research

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Vision foundation models are typically trained as static feature extractors, forcing the burden of task adaptation onto large downstream models. We propose a different paradigm: instead of solely feeding visual features into language, we use language itself to dynamically guide the vision encoder. Our method, Language-Instructed Vision Embeddings (LIVE), leverages language as high-level guidance to produce task-centric embeddings at inference time—without requiring task-specific retraining. This enables the encoder to focus attention on contextually relevant aspects of the input, yielding more controllable and generalizable representations. Empirically, LIVE reduces visual hallucinations (+34 points on MMVP), outperforms vision–language models with orders of magnitude more parameters on visual question answering, and generalizes to unseen instructions and tasks---offering a direct path toward adaptive, instruction-driven visual intelligence.