← 返回论文检索
IJCAI-ECAI 2026Main Track

Reexamining the Exploration–Exploitation Dilemma from an Entropy-Driven Perspective

Renye Yan, Yaozhong Gan, Jikang Cheng, Yi Sun, Zongwei Wang, Ling Liang, Yimao Cai

PDF 由论文原始站点提供,PaperCompass 不保存论文文件。

摘要

Achieving an optimal balance between exploration and exploitation remains a fundamental challenge in reinforcement learning. This work revisits the exploration-exploitation dilemma through the lens of entropy, offering a novel perspective on this enduring problem. It establishes a theoretical connection between policy's entropy and exploratory behavior, using entropy as a rational measure to quantify the exploration-exploitation trade-off. Theoretical analyses demonstrate that a modified Bellman equation, augmented with a novelty-seeking term, ensures appropriate entropy adjustment and guarantees its globally monotonic decay. The derived policy optimization process inherently accommodates all three regimes of the exploration-exploitation spectrum, enabling a principled transition from exploration to exploitation. Building on these theoretical insights, this work introduces AdaZero, an adaptive deep architecture that dynamically balances exploration and exploitation. Extensive empirical evaluations highlight AdaZero's robust performance and validate the theory's feasibility.