← 返回论文检索
ICML 2026PosterAccept (regular)

Proact-VL: A Proactive VideoLLM for Real-Time AI Companions

Weicai Yan, Yuhong Dai, Qi Ran, Haodong Li, Wang Lin, Hao Liao, Xing Xie, Tao Jin, Jianxun Lian

Zhejiang University,Tencent Hunyuan · StepFun · Shenzhen University · Department of Computer Science and Engineering, The Chinese University of Hong Kong · Zhejiang University · Microsoft Research Asia

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Proactive and real-time interactive experiences are essential for human-like AI companions, yet face three key challenges: (1) achieving low-latency inference under continuous streaming inputs, (2) autonomously deciding when to respond, and (3) controlling both quality and quantity of generated content to meet real-time constraints. In this work, we instantiate AI companions through two gaming scenarios—commentator and guide—selected for their suitability for automatic evaluation. We introduce the \textbf{Live Gaming Benchmark}, a large-scale dataset with three representative scenarios: solo commentary, co-commentary, and user guidance, and present \textbf{Proact-VL}, a general framework that shapes multimodal language models into proactive, real-time interactive agents capable of human-like environment perception and interaction. Extensive experiments show Proact-VL achieves superior response latency and quality while maintaining strong video understanding capabilities, demonstrating its practicality for real-time interactive applications. Code is available at https://anonymous.4open.science/r/Proact-VL-8699/.