כתבה
arXiv cs.LG ·
סגירת הפער האופקי באופטימיזציה של מדיניות
Closing the Horizon Gap in Policy Optimization for Adversarial MDPs
חוקרים פיתחו אלגוריתם חדש לאופטימיזציה של מדיניות בתהליכים מרקוביים עם הפסדים אדוורסריים. האלגוריתם משתמש בפונקציות Q מתוספות כדי לשלוט ביציבות של עדכונים מקומיים. התוצאות מראות שיפור בהתנהגות האופקית.
תקציר מקורי באנגליתarXiv:2610.12362v1 Announce Type: new Abstract: We consider policy optimization for online episodic tabular Markov decision processes (MDPs) with adversarial losses and bandit feedback. Policy optimization updates the policy locally at each state and avoids optimization over the occupancy-measure polytope, but its existing regret bounds are larger by a factor of the horizon $H$ than those of occupancy-measure-based algorithms. We close this gap by using regularized $Q$-functions, which allow us to control the stability of the local updates jointly over all state-action pairs rather than separately at each state. The resulting algorithm attains high-probability regret bounds of $\widetilde O(\sqrt{HS(H+A)T})$ for known transitions and $\widetilde O(HS\sqrt{AT})$ for unknown transitions, whe
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית