← 返回论文检索
IJCAI-ECAI 2026Main Track

BehaviorBench: A Psychologically Grounded Benchmark for Evaluating Personality in Large Language Models Through Realistic Behaviors

Taowen Pu, Hexi Wang, Zeyang Liu, Dongsheng Guo, Chuan Zhao

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Current approaches to evaluating personality in large language models (LLMs) typically prompt them to self-report on psychological questionnaires such as the Big Five Inventory. However, these methods assess introspective labels rather than observable behavior, despite the fact that LLMs are deployed to act in realistic contexts, not to reflect on their own traits. To bridge this gap, we introduce BehaviorBench, a new benchmark for evaluating LLM personality through concrete behaviors in everyday scenarios. Grounded in established personality psychology, BehaviorBench links each Big Five trait to validated behavioral manifestations and embeds them in contextually plausible situations that naturally elicit trait-relevant actions. We evaluate a range of models and personality shaping strategies using BehaviorBench and find a substantial mismatch between self-reported personalities and actual behaviors. By grounding evaluation in observable behavior rather than introspection, our work reveals critical gaps in current LLM personality modeling and control mechanisms. Our code and data are available at https://github.com/butra1n/BehaviorBench