NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
ByteDance Inc. · Peking University · Tianjin University · Beijing University of Aeronautics and Astronautics · Beihang University · Shenzhen University · ByteDance Seed · Beijing University of Posts and Telecommunications · China Jiliang University · University of Rochester · University of North Texas · The Hong Kong University of Science and Technology (Guangzhou) · Tsinghua University, Tsinghua University · Qihoo 360 · Tsinghua University · Abaka AI · Huazhong University of Science and Technology · nanjing university · The Chinese University of Hong Kong · Stanford University · Guangdong OPPO Mobile Telecommunications Corp.,Ltd. · Zhejiang University · Nanjing University · 2077AI · Tiktok · Alibaba Group
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. Experiments across state-of-the-art open- and closed-source models reveal that long-horizon repository generation remains largely unsolved, with even the strongest agents achieving merely 40\% average test pass rates and rarely completing an entire repository correctly. Further analysis identifies systematic long-horizon failure modes, including premature termination, loss of global coherence, fragile cross-file dependencies, and inadequate planning over hundreds of interaction steps. These results position NL2Repo-Bench as a rigorous, execution-based testbed for evaluating sustained agentic competence and highlight long-horizon reasoning as a key bottleneck for autonomous coding agents. Our data and code are available at https://anonymous.4open.science/r/nl2repobench-foricml-F4ED/.