← 返回论文检索
ICLR 2026PosterAccept (Poster)

Correlated Policy Optimization in Multi-Agent Subteams

Dingyang Chen, Jianing Ye, Zhenyu Zhang, Xiaolong Kuang, Xinyang Shen, Ozalp Ozer, Chongjie Zhang, Qi Zhang

Amazon · Institute for Interdisciplinary Information Sciences at Tsinghua University · Washington University in St. Louis · Worcester Polytechnic Institute

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

In cooperative multi-agent reinforcement learning, agents often face scalability challenges due to the exponential growth of the joint action and observation spaces. Inspired by the structure of human teams, we explore subteam-based coordination, where agents are partitioned into fully correlated subgroups with limited inter-group interaction. We formalize this structure using Bayesian networks and propose a class of correlated joint policies induced by directed acyclic graphs . Theoretically, we prove that regularized policy gradient ascent converges to near-optimal policies under a decomposability condition of the environment. Empirically, we introduce a heuristic for dynamically constructing context-aware subteams with limited dependency budgets, and demonstrate that our method outperforms standard baselines across multiple benchmark environments.