כתבה
arXiv cs.LG ·
יציבות מקומית וגלובלית בלמידת רפורמציה תפעולית
Local and Global Stability in Performative Reinforcement Learning
המאמר עוסק ביציבות בלמידת רפורמציה תפעולית, ומציע תיאוריה חדשה ליציבות זו. התיאוריה נותנת תנאים חדשים ליציבות, ומציעה דרך חדשה להבנת התהליך של למידת רפורמציה תפעולית.
תקציר מקורי באנגליתarXiv:2609.06467v1 Announce Type: new Abstract: In performative reinforcement learning the deployed policy shapes the environment that generates the learner's future data, and the natural solution concept is a performatively stable policy that is optimal in the environment it induces. Existing convergence guarantees rely on Lipschitz sensitivity assumptions on the environment map $\pi \mapsto (P_\pi, r_\pi)$, which are hard to verify and fail in settings such as multi-agent best-response dynamics. We instead study stability for mixtures of policies, and show that the resulting picture is fundamentally different from performative prediction, where randomization removes the need for any sensitivity assumption. We distinguish local mixed stability, an occupancy-weighted first-order relaxation
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית