Recursive Monte-Carlo Tree Search
Institute for Defense Analyses, Center for Communications Research, Princeton · Institute for Defense Analyses, Center for Communications Research, Princeton NJ
PDF 由论文原始站点提供,PaperCompass 不保存论文文件。
摘要
We introduce a recursive AlphaZero style Monte--Carlo tree search algorithm, "RMCTS". It first generates the search tree using prior policies, and then recursively re-estimates action values by using the regularized optimal posterior policies from ``Monte--Carlo tree search as regularized policy optimization'' (Grill et al., 2020) at each node of the search tree, starting from the leaves and working back up to the root. We find that RMCTS matches or exceeds the quality of AlphaZero's MCTS-UCB in a tiny fraction of the time.