CMPhysBench: A Benchmark for Evaluating Large Language Models in Condensed Matter Physics
Shanghai AI Laboratory · Institute of Physics, CAS · Hong Kong Polytechnic University · Tongji University · Institute of Physics Chinese Academy of Sciences · University of the Chinese Academy of Sciences · Fudan University · Shanghai Artificial Intelligence Laboratory · University of Hong Kong · Hunan Normal University · Lanzhou University of Technology · Zhengzhou University · University of Science and Technology of China · Kean College · Zhejiang University · Case Western Reserve University · Duke University · The Hong Kong University of Science and Technology (Guangzhou) · Shanghai AI Lab · Chinese Academy of Sciences · Hong Kong University of Science and Technology · Institute of Physics, Chinese Academy of Sciences
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
We introduce CMPhysBench, designed to assess the proficiency of Large Language Models (LLMs) in Condensed Matter Physics, as a novel Benchmark. CMPhysBench is composed of more than 520 graduate-level meticulously curated questions covering both representative subfields and foundational theoretical frameworks of condensed matter physics, such as magnetism, superconductivity, strongly correlated systems, etc. To ensure a deep understanding of the problem-solving process,we focus exclusively on calculation problems, requiring LLMs to independently generate comprehensive solutions. Meanwhile, leveraging tree-based representations of expressions, we introduce the Scalable Expression Edit Distance (SEED) score, which provides fine-grained (non-binary) partial credit and yields a more accurate assessment of similarity between prediction and ground-truth. Our results show that even the best models, Grok-4, reach only 36 average SEED score and 29% accuracy on CMPhysBench, underscoring a significant capability gap, especially for this practical and frontier domain relative to traditional physics.