יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

On the Approximation and Convergence of Distributional Policy Gradient Algorithms for Risk-Sensitive Reinforcement Learning

תקציר מקורי באנגליתarXiv:2405.14749v3 Announce Type: replace Abstract: Risk-sensitive reinforcement learning (RL) is crucial for maintaining reliable performance in high-stakes applications. While traditional RL methods aim to learn a point estimate of the random cumulative cost, distributional RL seeks to estimate the entire distribution of it, leading to a unified framework for handling different risk measures. However, developing policy gradient methods for risk-sensitive distributional RL is inherently more complex as it often involves finding the gradient of a probability measure. This paper introduces a new distributional policy gradient framework for risk-sensitive RL, where we derive an analytical gradient of the probability measure of the cumulative cost. For practical implementation, we further des
קרא במקור המקורי