← 返回论文检索
ICML 2026PosterAccept (regular)

Path-Coupled Bellman Flows for Distributional Reinforcement Learning

Boyang Xu, Qing Zou, Siqin Yang, Hao Yan

Arizona State University

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Distributional RL models the full return distribution, but common categorical/quantile approaches rely on projection and independently sampled Bellman targets, which ignore the Bellman operator’s affine transport structure and yield high-variance learning signals. We introduce Path-Coupled Bellman Flows, a flow-matching framework that shares base noise to couple the generative trajectories of consecutive states, inducing a geometric Bellman scaling law between their velocity fields. This geometry motivates a $\lambda$-family of Bellman-flow objectives that functions as a control variate, reducing variance while retaining the same Bellman-consistent fixed point. Across toy diagnostics and offline RL benchmarks (OGBench, D4RL), our method improves training stability and achieves competitive or improved performance relative to prior distributional baselines.