Identifying Latent State-Transition Processes for Individualized Reinforcement Learning
Mohamed bin Zayed University of Artificial Intelligence · University of California, San Diego · University of Sydney · KDDI Corporation · Carnegie Mellon University · KDDI Research, Inc. · Osaka University · CMU & MBZUAI
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
The application of reinforcement learning (RL) involving interactions with individuals has grown significantly in recent years. These interactions, influenced by factors such as personal preferences and physiological differences, causally influence state transitions, ranging from health conditions in healthcare to learning progress in education. As a result, different individuals may exhibit different state-transition processes. Understanding individualized state-transition processes is essential for optimizing individualized policies. In practice, however, identifying these state-transition processes is challenging, as individual-specific factors often remain latent. In this paper, we establish the identifiability of these latent factors and introduce a practical method that effectively learns these processes from observed state-action trajectories. Experiments on various datasets show that the proposed method can effectively identify latent state-transition processes and facilitate the learning of individualized RL policies.