← 返回论文检索
ICLR 2026PosterAccept (Poster)

Negotiated Reasoning: On Provably Addressing Relative Over-Generalization

Junjie Sheng, Yantian Wang, Bo Jin, Hongyuan Zha, Jun Wang, Wenhao Li, Xiangfeng Wang

East China Normal University · Tongji University · College of Computing, Georgia Institute of Technology · East China Normal University, Tsinghua University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

We focus on the relative over-generalization (RO) issue in fully cooperative multi-agent reinforcement learning (MARL). Existing methods show that endowing agents with reasoning can help mitigate RO empirically, but there is little theoretical insight. We first prove that RO is avoided when agents satisfy a consistent reasoning requirement. We then propose a new negotiated reasoning framework connecting reasoning and RO with theoretical guarantees. Based on it, we develop an algorithm called Stein variational negotiated reasoning (SVNR), which uses Stein variational gradient descent to form a negotiation policy that provably bypasses RO under maximum-entropy policy iteration. SVNR is further parameterized with neural networks for computational efficiency. Experiments demonstrate that SVNR significantly outperforms baselines on RO-challenged tasks, including Multi-Agent Particle World and MaMuJoCo, confirming its advantage in achieving better cooperation.