Visual and Memory–Augmented Soccer Commentary Generation
Japan Advanced Institute of Science and Technology, Tokyo Institute of Technology
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-long.485 ↗
摘要
Automatic soccer commentary generation aims to bridge the gap between raw visual content and professional, tactical commentary. However, existing datasets tend to produce incomplete commentary that lacks semantic richness and fails to convey the full visual information present in standard video clips. To address these limitations, we propose two manually curated datasets: SN-Short, which enhances scene-level semantic descriptions, and SN-Long, which captures event continuity for context-aware commentary.Based on these, we design a commentary augmentation pipeline that transforms incomplete annotations into MatchText, a semantically complete and structurally standardized dataset. Leveraging this supervision, we introduce MatchAware, a generation model that incorporates contextual cues from previous events to produce coherent commentary aligned with the visual flow of the game. Experimental results show that proposed approach significantly outperforms existing baselines on the constructed datasets.