MTRAG-UN: A Benchmark for Open Challenges in Multi-Turn RAG Conversations
International Business Machines
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2026.findings-acl.503 ↗
摘要
We present MTRAG-UN, a benchmark for exploring open challenges in multi-turn retrieval augment generation, a popular use of large language models. We release a benchmark of 666 tasks from 666 conversations containing over 2,800 conversation turns across 6 domains with accompanying corpora. Our experiments show that retrieval and generation models continue to struggle on conversations with UNanswerable, UNderspecified, and NONstandalone questions and UNclear responses. Our benchmark is available at https://github.com/IBM/mt-rag-benchmark