CentaurEval: Benchmarking Human-in-the-Loop Value in Agentic Coding
New York University Abu Dhabi · University of Illinois at Urbana-Champaign · Microsoft · Beijing University of Technology · University of Electronic Science and Technology of China · Zhejiang University · Hong Kong Polytechnic University · New York University · Nanyang Technological University · Indiana University
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
LLM-powered coding agents are reshaping the development paradigm. However, existing evaluation systems, neither traditional tests for humans nor benchmarks for LLMs, fail to capture this shift, excluding problems that require both human reasoning to guide solutions and AI efficiency for implementation. We introduce CentaurEval, a unified, ecologically valid benchmark for measuring human-in-the-loop value in coding. CentaurEval's core innovation is its "Collaboration-Necessary" problem templates, which are intractable for standalone LLMs or humans, but solvable through effective collaboration. CentaurEval dynamically instantiates tasks from 45 templates, providing a standardized IDE for humans and a reproducible 450-task toolkit for LLMs. We benchmark 45 participants against 5 LLMs under 4 levels of human intervention. Results show that while LLMs or humans alone achieve poor pass rates (0.67% and 18.89%), human–AI collaboration significantly improves to 31.11%. Our analysis reveals an emerging co-reasoning partnership, challenging the traditional human-tool hierarchy by showing that strategic breakthroughs can originate from either humans or AI. Our work is openly accessible.