כתבה
arXiv cs.LG ·
FERPO: Forward Entropy-Regularized Policy Optimization
תקציר מקורי באנגליתarXiv:2610.02198v1 Announce Type: new Abstract: Several state-of-the-art methods for online reinforcement learning in continuous control improve policies using action gradients of a learned critic. However, critics are typically trained to predict returns, and accurate value predictions do not necessarily yield accurate action derivatives, potentially leading to unreliable policy updates. We propose Forward Entropy-Regularized Policy Optimization (FERPO), an on-policy maximum entropy reinforcement learning algorithm that performs policy improvement using critic values without differentiating the critic with respect to actions. FERPO derives an optimal target action distribution from a policy-improvement objective regularized by entropy and Kullback-Leibler (KL) divergence. We then fit the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית