← 返回论文检索
ACL 2026shortmain

GOLEMcoref: A Multilingual Coreference Dataset of Fiction

Andreas van Cranenburgh, Xiaoyan Yang, Alvanita, Cecilia Nicole Di Domenico, Maria Ferragud, Arianna Graciotti, Andreea Gabriela Ion, Byungjun Kim, Seonyeong Park, Noa Visser Solissa, Xiaoyu Zhou, Federico Pianzola

University of Groningen · Coventry University and Universitas Gadjah Mada · University of Groningen, Fondazione Bruno Kessler and Università di Trento · The Academy of Korean Studies

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.acl-short.39 ↗

摘要

We present a multilingual coreference dataset of 827k tokens of fiction in 7 languages: Bahasa Indonesia, Chinese, Dutch, English, Italian, Korean, and Spanish. The dataset includes full stories of diverse lengths, ranging from 500 to 17k words. We discuss our annotation scheme focusing on characters and language-specific challenges we encountered. Finally we present evaluation results of a neural coreference system trained on our dataset. We show that jointly training a system across all languages provides a strong improvement over monolingually trained models. The dataset is available under a creative commons license in CoNLL-2012 and CorefUD format at https://github.com/GOLEM-lab/GOLEMcoref/