← 返回论文检索
NeurIPS 2025{location} PosterAccept (poster)

SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Xeron Du, Yifan Yao, Kaijing Ma, Bingli Wang, Tianyu Zheng, Zhu, Minghao Liu, Yiming Liang, Xiaolong Jin, Zhenlin Wei, Chujie Zheng, Kaixin Deng, Shuyue Guo, Shian Jia, Sichao Jiang, Yiyan Liao, Rui Li, Qinrui Li, Sirun Li, Yizhi Li, Yunwen Li, Dehua Ma, Yuansheng Ni, Haoran Que, Qiyao Wang, Zhoufutu Wen, Siwei Wu, Tianshun Xing, 明 许, Zhenzhu Yang, Noah Wang, Junting Zhou, yuelin bai, Xingyuan Bu, chenglin cai, Liang Chen, Yifan Chen, Cheng Chengtuo, Tianhao Cheng, Keyi Ding, Siming Huang, HUANG YUN, Yaoru Li, Yizhe Li, Zhaoqun Li, Tianhao Liang, Chengdong Lin, Hongquan Lin, Yinghao Ma, Zhongyuan Peng, Zifan Peng, Qige Qi, Shi Qiu, Xingwei Qu, Shanghaoran Quan, Yizhou Tan, Zili Wang, 王晨清, Hao Wang, Yiya Wang, Yubo Wang, Jiajun Xu, Kexin Yang, Ruibin Yuan, Yuanhao Yue, Tianyang Zhan, Chun Zhang, Jinyang Zhang, Xiyue Zhang, Owen Zhang, Yue Zhang, Yongchi Zhao, Xiangyu Zheng, ChenghuaZhong, Yang Gao, Zhoujun Li, Dayiheng Liu, Qian Liu, Tianyu Liu, Shiwen Ni, Junran Peng, Yujia Qin, Wenbo Su, Guoyin Wang, Shi Wang, Jian Yang, Min Yang, Meng Cao, Xiang Yue, ZHAO-XIANG ZHANG, Wangchunshu Zhou, Jiaheng Liu, Qunshu Lin, Wenhao Huang, Ge Zhang

01.AI · Beijing University of Posts and Telecommunications · Tongji University · Sichuan Agricultural University · Guangdong OPPO Mobile Telecommunications Corp.,Ltd. · 2077AI · University of the Chinese Academy of Sciences · Purdue University · Harbin Engineering University · Tsinghua University · Hokkaido University · Zhejiang University · zhejiang university · Peking University · Cornell University · The University of Manchester · Chinese University of Hong Kong(shenzhen) · University of Waterloo · Beijing University of Aeronautics and Astronautics · henzhen Institute of Advanced Technology, Chinese Academy of Sciences · ByteDance Inc. · Nanjing University of Science and Technology · China University of Geoscience Beijing · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences, Chinese Academy of Sciences · Alibaba Group · Huawei Technologies Ltd. · Fudan University · University of Melbourne · national university of singaore, National University of Singapore · Hangzhou Dianzi University · University of Science and Technology of China · Centre for Digital Music, Queen Mary University of London · The Hong Kong University of Science and Technology (Guangzhou) · University of Manchester · Harvard University · stepfun · abaka · Facebook · Carnegie Mellon University · Department of Computer Science, Princeton University · Suzhou University · University of Science and Technology Beijing · Nanjing University · TikTok (Singapore) · Alibaba · Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences · Institute of automation, Chinese academy of science · Bytedance · Alibaba Qwen Pilot · Institute of Computing Science, Chinese Academy of Sciences · Mohamed bin Zayed University of Artificial Intelligence · Meta · Chinese Academy of Sciences, China · Abaka AI · Key Laboratory of Machine Perception · University of Michigan - Ann Arbor

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Large language models (LLMs) have demonstrated remarkable proficiency in mainstream academic disciplines such as mathematics, physics, and computer science. However, human knowledge encompasses over 200 specialized disciplines, far exceeding the scope of existing benchmarks. The capabilities of LLMs in many of these specialized fields-particularly in light industry, agriculture, and service-oriented disciplines-remain inadequately evaluated. To address this gap, we present SuperGPQA, a comprehensive benchmark that evaluates graduate-level knowledge and reasoning capabilities across 285 disciplines. Our benchmark employs a novel Human-LLM collaborative filtering mechanism to eliminate trivial or ambiguous questions through iterative refinement based on both LLM responses and expert feedback. Our experimental results reveal significant room for improvement in the performance of current state-of-the-art LLMs across diverse knowledge domains (e.g., the reasoning-focused model Gemini-2.5-Pro achieved the highest accuracy of 63.56% on SuperGPQA), highlighting the considerable gap between current model capabilities and artificial general intelligence. Additionally, we present comprehensive insights from our management of a large-scale annotation process, involving over 80 expert annotators and an interactive Human-LLM collaborative system, offering valuable methodological guidance for future research initiatives of comparable scope.