KoLA: Carefully Benchmarking World Knowledge of Large Language Models
Tsinghua University, Tsinghua University · Tsinghua University · Zhipu AI · Beijing Knowledge Atlas Technology Co., Ltd. · Department of Computer Science and Technology, Tsinghua University, Tsinghua University · Beijing University of Posts and Telecommunications · DCST, Tsinghua University · Department of Computer Science and Technology, Tsinghua University · , Tsinghua University · University of California, Los Angeles, Computer Science Department · Ohio State University · Department of Computer Science, Tsinghua University
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
The unprecedented performance of large language models (LLMs) necessitates improvements in evaluations. Rather than merely exploring the breadth of LLM abilities, we believe meticulous and thoughtful designs are essential to thorough, unbiased, and applicable evaluations. Given the importance of world knowledge to LLMs, we construct a Knowledge-oriented LLM Assessment benchmark (KoLA), in which we carefully design three crucial factors: (1) For ability modeling, we mimic human cognition to form a four-level taxonomy of knowledge-related abilities, covering 19 tasks. (2) For data, to ensure fair comparisons, we use both Wikipedia, a corpus prevalently pre-trained by LLMs, along with continuously collected emerging corpora, aiming to evaluate the capacity to handle unseen data and evolving knowledge. (3) For evaluation criteria, we adopt a contrastive system, including overall standard scores for better numerical comparability across tasks and models, and a unique self-contrast metric for automatically evaluating knowledge-creating ability. We evaluate 21 open-source and commercial LLMs and obtain some intriguing findings. The KoLA dataset will be updated every three months to provide timely references for developing LLMs and knowledge-related systems.