← 返回论文检索
ICML 2026PosterAccept (spotlight)

Recurrent Structural Policy Gradient for Partially Observable Mean Field Games

Clarisse Wibault, Sebastian Towers, Tiphaine Wibault, Juan Duque, Johannes Forkel, George Whittle, Andreas Schaab, Chiyuan Wang, Yucheng Yang, Michael A Osborne, Benjamin Moll, Jakob Foerster

University of Oxford · University of Oxford, University of Oxford · Ludwig-Maximilians-Universität München · Mila Quebec AI Institute · University of California, Berkeley · Peking University · University of Zurich · U Oxford · London School of Economics and Political Science, University of London · Oxford university

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Mean Field Games (MFGs) provide a principled framework for modeling interactions in large populations models: at scale, population dynamics become deterministic, with uncertainty entering only through aggregate shocks, or *common noise*. However, algorithmic progress has been limited since model-free methods are too high variance and exact methods scale poorly. Recent Hybrid Structural Methods (HSMs) use Monte Carlo rollouts for the common noise in combination with exact estimation of the expected return, conditioned on those samples. However, HSMs have not been scaled to Partially Observable settings. We propose *Recurrent Structural Policy Gradient* (RSPG), the first history-aware HSM. We also introduce MFAX, our JAX-based framework for MFGs. By leveraging known transition dynamics, RSPG achieves state-of-the-art performance as well as an order-of-magnitude faster convergence and solves, for the first time, a macroeconomics MFG with heterogeneous agents, common noise and history-aware policies. MFAX is publicly available at: .