Video Question Answering and Beyond
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.1145/3746027.3759207 ↗
摘要
Video Question Answering (Video QA) has emerged as a central task in multimodal learning. This tutorial provides a comprehensive overview of VideoQA research and highlights new frontiers. We begin with an introduction to VideoQA preliminaries, tracing how methods have adapted from third-person view short videos to capture egocentric and long-ranged spatial-temporal dynamics. We then focus on the impact of large multimodal models. Next, we expand the scope to spatial understanding beyond videos. Each topic is discussed through the lens of tasks, datasets, methods, and evaluation protocols. Finally, we conclude with future directions, including fine-grained and long-ranged video understanding, robustness and trustworthiness, Egocentric and embodied assistance, and omnimodal integration. This tutorial aims to equip participants with both a historical perspective and a forward-looking roadmap for advancing Video QA in the LLM era.