← 返回论文检索
ICML 2026PosterAccept (regular)

Position: Time to Close The Validation Gap in LLM Social Simulations

Maximilian Puelma Touzel, Sneheel Sarangi, Aurélien Bück-Kaeffer, Zachary Yang, Jean-François Godbout, Reihaneh Rabbany

MILA - Quebec AI Institute · McGill University · McGill University, Mila, Ubisoft La Forge · McGill | Mila | Ubisoft La Forge · Université de Montréal · McGill University, Mila

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

LLM-based social simulations—in which many language model agents interact over multiple turns—are rapidly proliferating across policy analysis, epidemiology, and computational social science. Yet the field lacks consensus on how to validate these simulations, with evaluation methods that are sparse, inconsistent, and rarely shared across disciplinary silos. We argue this creates a serious risk: premature deployment of unvalidated simulators in high-stakes domains. Our position is that the field must pivot from expansion to consolidation, prioritizing methodological standardization—shared benchmarks, open data, and reproducible evaluation protocols grounded in social science and complex systems research. We outline a concrete research program organized around specific learning problems/benchmarks, providing a path toward answering the fundamental question: when are LLM social simulations useful modelling objects?