← 返回论文检索
ICML 2026PosterAccept (regular)

NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents

Jingzhe Ding, Shengda Long, Changxin Pu, Ge Zhang, zhou huan, Hongwan Gao, Xiang Gao, Chao He, Yue Hou, FEI HU, Zhaojian Li, Weiran Shi, Zaiyuan Wang, Daoguang Zan, Chenchen Zhang, Xiaoxu Zhang, Chen Qizhi, cheng, Bo Deng, Qingshui Gu, Kai Hua, Juntao Lin, Pai Liu, Mingchen Li, Minghao Li, Xuanguang Pan, Zifan Peng, Yujia Qin, Yong Shan, Zhewen Tan, Haoran Wang, Tom Tang, Weihao Xie, Yishuo Yuan, Jiayu Zhang, Yunfei Zhao, He Zhu, LIYA ZHU, chenyangzou, Ming Ding, Jiaheng Liu, Jianpeng Jiao, Liam Liu, Qian Liu, Chongyang Tao, Jian Yang, Tong Yang, Zhaoxiang Zhang, Xinjie Chen, Wenhao Huang

ByteDance Inc. · Peking University · Tianjin University · Beijing University of Aeronautics and Astronautics · Beihang University · Shenzhen University · ByteDance Seed · Beijing University of Posts and Telecommunications · China Jiliang University · University of Rochester · University of North Texas · The Hong Kong University of Science and Technology (Guangzhou) · Tsinghua University, Tsinghua University · Qihoo 360 · Tsinghua University · Abaka AI · Huazhong University of Science and Technology · nanjing university · The Chinese University of Hong Kong · Stanford University · Guangdong OPPO Mobile Telecommunications Corp.,Ltd. · Zhejiang University · Nanjing University · 2077AI · Tiktok · Alibaba Group

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Recent advances in coding agents suggest rapid progress toward autonomous software development, yet existing benchmarks primarily evaluate short-horizon behaviors such as localized code generation, scaffolded completion, or repository repair, leaving it unclear whether agents can sustain coherent reasoning, planning, and execution over the extended horizons demanded by real-world repository construction. To address this gap, we introduce NL2Repo-Bench, a benchmark explicitly designed to evaluate the long-horizon repository generation from scratch: given only a single natural-language requirements document and an empty workspace, agents must autonomously design the architecture, manage dependencies, and produce a fully installable Python library. Experiments across state-of-the-art open- and closed-source models reveal that long-horizon repository generation remains largely unsolved, with even the strongest agents achieving merely 40\% average test pass rates and rarely completing an entire repository correctly. Further analysis identifies systematic long-horizon failure modes, including premature termination, loss of global coherence, fragile cross-file dependencies, and inadequate planning over hundreds of interaction steps. These results position NL2Repo-Bench as a rigorous, execution-based testbed for evaluating sustained agentic competence and highlight long-horizon reasoning as a key bottleneck for autonomous coding agents. Our data and code are available at https://anonymous.4open.science/r/nl2repobench-foricml-F4ED/.