GameVerse: Can Vision-Language Models Learn from Video-based Reflection?
Tsinghua University · Chongqing University · Beijing Institute of Technology · Dalian University of Technology · NVIDIA
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Human gameplay is a visually grounded interaction loop in which players act, reflect on failures, and watch tutorials to refine strategies. Can Vision-Language Models (VLMs) also learn from video-based reflection? We present **GameVerse**, a comprehensive video game benchmark that enables a *reflective visual interaction loop*. Moving beyond traditional ***fire-and-forget*** evaluations, it uses a novel ***reflect-and-retry*** paradigm to assess how VLMs internalize visual experience and improve policies. To facilitate systematic and scalable evaluation, we also introduce a *cognitive hierarchical taxonomy* spanning 15 globally popular games, *dual action space* for both semantic and GUI control, and *milestone evaluation* using advanced VLMs to quantify progress. Our experiments show that VLMs benefit from video-based reflection in varied settings, and perform best by combining failure trajectories and expert tutorials—a *training-free* analogue to reinforcement learning (RL) plus supervised fine-tuning (SFT).