← 返回论文检索
ACL 2024aclfindings

PUB: A Pragmatics Understanding Benchmark for Assessing LLMs’ Pragmatics Capabilities

Settaluri Sravanthi, Meet Doshi, Pavan Tankala, Rudra Murthy, Raj Dabre, Pushpak Bhattacharyya

Indian Institute of Technology Bombay, Indian Institute of Technology, Bombay · IBM India Ltd · National Institute of Information and Communications Technology (NICT), National Institute of Advanced Industrial Science and Technology · Indian Institute of Technology, Bombay, Dhirubhai Ambani Institute Of Information and Communication Technology

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2024.findings-acl.719 ↗

摘要

LLMs have demonstrated remarkable capability for understanding semantics, but their understanding of pragmatics is not well studied. To this end, we release a Pragmatics Understanding Benchmark (PUB) dataset consisting of fourteen tasks in four pragmatics phenomena, namely; Implicature, Presupposition, Reference, and Deixis. We curate high-quality test sets for each task, consisting of Multiple Choice Question Answers (MCQA). PUB includes a total of 28k data points, 6.1k are newly annotated. We evaluate nine models varying in the number of parameters and type of training. Our study reveals several key observations about the pragmatic capabilities of LLMs: 1. chat-fine-tuning strongly benefits smaller models, 2. large base models are competitive with their chat-fine-tuned counterparts, 3. there is a huge variance in performance across different pragmatics phenomena, and 4. a noticeable performance gap between human capabilities and model capabilities. We hope that PUB will enable comprehensive evaluation of LLM’s pragmatic reasoning capabilities.