← 返回论文检索
ICLR 2026Blog Track PosterAccept (Poster)

From REINFORCE to Dr. GRPO: A Unified Perspective on LLM Post-Training

Qingfeng Lan

University of Alberta

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Recently, many reinforcement learning (RL) algorithms have been applied to improve the post-training of large language models (LLMs). In this article, we aim to provide a unified perspective on the objectives of these RL algorithms, exploring how they relate to each other through the Policy Gradient Theorem — the fundamental theorem of policy gradient methods.