SWE-Compass: Towards Unified Evaluation of Agentic Coding Abilities for Large Language Models
Kuaishou Technology · Kuaishou- 快手科技 · Southern University of Science and Technology · ByteDance Inc. · Beijing University of Posts and Telecommunications · BUPT · China Telecom · Nanjing University · nanjing university · Dalian University of Technology · Hong Kong Polytechnic University · Li Auto Inc. · National University of Singapore · Beijing University of Post and Telecommunications · Xiaomi Corporation · Minzu University of China · Singapore University of Technology and Design · Kuaishou Tech
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Evaluating large language models (LLMs) for software engineering has been limited by narrow task coverage, language bias, and insufficient alignment with real-world developer workflows. Existing benchmarks often focus on algorithmic problems or Python-centric bug fixing, leaving critical dimensions of software engineering underexplored. To address these gaps, we introduce SWE-Compass, a comprehensive benchmark that unifies heterogeneous code-related evaluations into a structured and production-aligned framework. SWE-Compass spans 8 task types, 8 programming scenarios, and 10 programming languages, with 2000 high-quality instances curated from authentic GitHub pull requests and refined through systematic filtering and validation. We benchmark ten state-of-the-art LLMs under two agentic frameworks, SWE-Agent and Claude Code, revealing a clear hierarchy of difficulty across task types, languages, and scenarios. Moreover, by aligning evaluation with real-world developer practices, we hope SWE-Compass can provide a rigorous and reproducible foundation for diagnosing and advancing agentic coding capabilities in large language models.