Adaptively profiling models with task elicitation
University of Pennsylvania, University of Pennsylvania and Pacific Northwest National Laboratory · University of Pennsylvania, University of Pennsylvania · University of Pennsylvania · University of Pennsylvania, University of Pennsylvania and University of Pennsylvania
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。DOI 10.18653/v1/2025.emnlp-main.1270 ↗
摘要
Language model evaluations often fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks. We introduce task elicitation, a method that automatically builds new evaluations to profile model behavior. Task elicitation finds hundreds of natural-language tasks—an order of magnitude more than prior work—where frontier models exhibit systematic failures, in domains ranging from forecasting to online harassment. For example, we find that Sonnet 3.5 over-associates quantum computing and AGI and that o3-mini is prone to hallucination when fabrications are repeated in-context.