QEDBench: Quantifying the Alignment Gap in Automated Evaluation of University-Level Mathematical Proofs
Yale University · Sharif University of Technology · University of California, Berkeley · Charles University Prague · University of Chicago · University of Illinois at Urbana-Champaign · MIT · University of Delaware · Princeton University · Massachusetts Institute of Technology · Universita della Svizzera Italiana · Caltech · Technische Universität Hamburg · Universität Regensburg · University of Helsinki · Illinois Institute of Technology · University of Western Ontario · University of Georgia · Fort Lewis College · Helm.ai · University of British Columbia · CUNY City Tech · Universität Köln · University of Amsterdam · King's College London · UPF & Google DeepMind · The University of Edinburgh · Bilkent University · York University · University of Waterloo · Ahmedabad University · Burapha University · Johannes-Gutenberg Universität Mainz · CUNY Graduate Center · Nagoya University · University of Oxford · Independent · Bursa Uludag University · Alfréd Rényi Institute of Mathematics
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
As Large Language Models (LLMs) saturate elementary benchmarks, the research frontier has shifted from generation to the reliability of automated evaluation. We demonstrate that standard "LLM-as-a-Judge" protocols suffer from a systematic evaluation Alignment Gap when applied to upper-undergraduate to early graduate level mathematics. To quantify this, we introduce QEDBench, the first benchmark to systematically measure alignment with human experts on undergraduate-level math proofs by contrasting course-specific rubrics against expert common knowledge criteria. By deploying a dual-evaluation matrix ($7$ judges $\times$ $5$ solvers) against 1,000+ hours of human evaluation, we reveal that certain frontier evaluators like Claude 4.5 Opus exhibit significant positive bias (up to $+0.28$ mean score inflation), effectively "hallucinating rigor" in flawed proofs. Furthermore, we uncover a critical reasoning disparity: while Gemini 3.0 Pro achieves state-of-the-art performance (0.91 raw score), specialized reasoning models like o3-deep-research collapse in discrete domains, dropping to 42.1\% accuracy in Graph Theory. We release QEDBench as a public benchmark for evaluating and improving AI judges.