Learning to Maximize Rewards via Reaching Goals
Princeton University · Carnegie Mellon University
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
Goal-conditioned reinforcement learning learns to reach goals instead of optimizing hand-crafted rewards. Despite its popularity, the community often categorizes goal-conditioned reinforcement learning as a special case of reinforcement learning. In this post, we aim to build a direct conversion from any reward-maximization reinforcement learning problem to a goal-conditioned reinforcement learning problem, and to draw connections with the stochastic shortest path framework. Our conversion provides a new perspective on the reinforcement learning problem: *maximizing rewards is equivalent to reaching some goals*.