Principled RL for Flow Matching Emerges From the Chunk-level Policy Optimization
Tsinghua University · Beihang University · Kuaishou- 快手科技 · Zhejiang University · Tsinghua University, Tsinghua University · Shenzhen International Graduate School, Tsinghua University · Alibaba Group
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Recent Progress in post-training flow matching for text-to-image (T2I) generation with Group Relative Policy Optimization (GRPO) has demonstrated strong potential. However, it is hindered by a critical limitation: inaccurate advantage attribution. In this work, we argue that aggregating consecutive timesteps into a coherent `chunk' and shifting the policy optimization paradigm from GRPO's step level to the chunk level can effectively mitigate the negative impact of this issue. Building on this insight, we propose Group Chunking Policy Optimization (GCPO), the first chunk-level reinforcement learning approach for post-training flow matching. Extensive experiments demonstrate that GCPO achieves superior performance on both standard T2I benchmarks and preference alignment, with up to $43\%$ additional gains over GRPO, highlighting the promise of chunk-level policy optimization.