← 返回论文检索
IJCAI-ECAI 2026Sister Conferences Best Papers Track

Soft Condorcet Optimization for Ranking of General Agents (Extended Abstract)

Marc Lanctot, Kare Larson, Michael Kaisers, Quentin Berthet, Ian Gemp, Manfred Diaz, Roberto-Rafael Maura-Rivero, Yoram Bachrach, Anna Koop, Doina Precup

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Driving progress of AI models and agents requires comparing their performance on standardized benchmarks; for general agents, individual performances must be aggregated across a potentially wide variety of different tasks. In this extended abstract, we describe a ranking scheme inspired by social choice frameworks, called Soft Condorcet Optimization (SCO), to compute the optimal ranking of agents: the one that makes the fewest mistakes in predicting the agent comparisons in the evaluation data. This optimal ranking is the maximum likelihood estimate when evaluation data (which we view as votes) are interpreted as noisy samples from a ground truth ranking, a solution to Condorcet's original voting system criteria. SCO ratings are maximal for Condorcet winners when they exist, which we show is not necessarily true for the classical rating system Elo. In practice, SCO serves as an accurate approximation to the Kemeny-Young voting method, excels in the sparse data regime, and provides the best approximation to the optimal ranking compared to every baseline on a Diplomacy player ranking problem with 31,094 games and 52,958 agents.