כתבה
arXiv cs.LG ·
התכנסות של תנאי הצלחה לאופטימיזציה של מדיניות
On the Convergence of Success Conditioning for Policy Optimization
המאמר חושף את התכנסות של תנאי הצלחה לאופטימיזציה של מדיניות בתהליכי החלטה שוטוריים. נושא זה חשוב ללמידת מכונה ולשיפור רשתות עצבים.
תקציר מקורי באנגליתarXiv:2610.03642v1 Announce Type: new Abstract: Success conditioning is a strategy for improving decision-making policies in stochastic environments; it updates a policy by increasing the probability of taking actions that yield successful outcomes. Success conditioning is common to many reinforcement learning applications, yet its limiting behavior and convergence rates are not well understood. In this work, we demonstrate that success conditioning converges to an optimal policy on a broad class of Markov decision processes (MDPs). We also derive convergence rates in some common settings. For discounted MDPs, we prove convergence within $\mathcal{O}(1/\varepsilon^p)$ iterations to an $\varepsilon$-optimal policy, where the exponent $p$ depends on problem data. For single-period MDPs, such
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית