← 返回论文检索
ICML 2026PosterAccept (regular)

Bridging the Gap in Autonomous Science: The Corpus and Benchmark for Biological Protocol Reasoning

Yuyang Liu, Liuzhenghao Lyu, Xiancheng Zhang, Jingya Wang, Li Yuan, Yonghong Tian

Peking University · Harbin Institute of Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

The realization of autonomous scientific experimentation is currently limited by LLMs' struggle to grasp the strict procedural logic and accuracy required by biological protocols. To address this fundamental challenge, we present **BioProBench**, a comprehensive resource for procedural reasoning in biology. BioProBench is grounded in**BioProCorpus**, a foundational collection of 27,000 human-written protocols. From this corpus, we systematically constructed a dataset of over 550,000 task instances, offering both a large-scale training resource and a rigorous benchmark with novel metrics. Evaluating 10 mainstream LLMs, we find that while general comprehension is high, performance drops significantly on tasks demanding deep reasoning, quantitative precision, and safety awareness. To demonstrate the value of BioProCorpus in mitigating these issues, we developed **ProAgent**, grounded in our corpus, ProAgent substantially advances the state-of-the-art. BioProBench provides a rigorous diagnostic benchmark and a foundational resource for developing the next generation of reliable scientific AI. Code and data are available at: https://anonymous.4open.science/r/Anonymization-112358 .