כתבה
arXiv cs.LG ·
חשיבה מחדש על גלובליות סופטמקס פוליצי גרדיאנט עם תיאור פונקציונלי לינארי: המקרה של חביות חד-זרוע
Rethinking the Global Convergence of Softmax Policy Gradient with Linear Function Approximation: The Case of Multi-Armed Bandits
במאמר זה, חוקרים חושבים מחדש על גלובליות סופטמקס פוליצי גרדיאנט עם תיאור פונקציונלי לינארי, ומראים שהטעות באפוקסימציה אינה חשובה לגלובליות האלגוריתם, גם במקרה של חביות חד-זרוע.
תקציר מקורי באנגליתarXiv:2505.03155v2 Announce Type: replace Abstract: Policy gradient (PG) methods have played an essential role in the empirical successes of reinforcement learning. In order to handle large state-action spaces, PG methods are typically used with function approximation. In this setting, the approximation error in modeling problem-dependent quantities is a key notion for characterizing the global convergence of PG methods. We study Softmax PG with linear function approximation (referred to as $\texttt{Lin-SPG}$) and demonstrate that the approximation error is irrelevant to the algorithm's global convergence even in the bandit setting. Consequently, we rethink the effect of approximation error in the standard stochastic multi-armed bandit problem. We first identify the conditions on the polic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית