← 返回论文检索
EMNLP 2024mainmain

LawBench: Benchmarking Legal Knowledge of Large Language Models

Zhiwei Fei, Xiaoyu Shen, Dawei Zhu, Fengzhe Zhou, Zhuo Han, Alan Huang, Songyang Zhang, Kai Chen, Zhixin Yin, Zongwen Shen, Jidong Ge, Vincent Ng

Fudan University, Harbin Institute of Technology, Dalian University of Technology, Shanghai Jiaotong University, Shandong University, Peking University, Zhejiang University, University of Science and Technology of China, Hunan University, Beijing Institute of Technology, University of the Chinese Academy of Sciences, Southeast University, Sichuan University, Monash University, Malaysia Campus, Tianjin University, Beijing University of Aeronautics and Astronautics, Wuhan University of Technology, Yale University, Technische Universität München, Wuhan University, nanjing university, Tsinghua University and Wuhan University · Amazon · Science and Engineering Magnet School · Shanghai AI Laboratory · Nanjing University · University of Texas at Dallas

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2024.emnlp-main.452 ↗

摘要

We present LawBench, the first evaluation benchmark composed of 20 tasks aimed to assess the ability of Large Language Models (LLMs) to perform Chinese legal-related tasks. LawBench is meticulously crafted to enable precise assessment of LLMs’ legal capabilities from three cognitive levels that correspond to the widely accepted Bloom’s cognitive taxonomy. Using LawBench, we present a comprehensive evaluation of 21 popular LLMs and the first comparative analysis of the empirical results in order to reveal their relative strengths and weaknesses. All data, model predictions and evaluation code are accessible from https://github.com/open-compass/LawBench.