כתבה
arXiv cs.LG ·
התכנסות בזמן סופי של למידת Q רובוסטית
Finite-Time Convergence of Single-Trajectory Chi-Square Robust Q-Learning With Linear Function Approximation
חוקרים פיתחו שיטה חדשה ללמידת Q רובוסטית עם אי-ודאות כי-רבוע. השיטה משתמשת בניסוח וריאציוני ובסכמת יעד קפוא. המחקר מראה תוצאות מבטיחות במשימות בקרה לא-ליניאריות.
תקציר מקורי באנגליתarXiv:2510.01721v4 Announce Type: replace Abstract: Distributionally robust reinforcement learning seeks policies that remain effective when the deployment environment differs from the one that generated the training data. We study model-free robust Q-learning with $\chi^2$ uncertainty sets and linear function approximation, using data from a single trajectory of an unknown nominal MDP. Evaluating the $\chi^2$ robust Bellman target introduces the square root of a conditional second moment, which cannot be estimated unbiasedly from one transition, while the projected robust Bellman operator need not be contractive. We address these obstacles through a variational reformulation of the robust Bellman target and a blockwise frozen-target scheme, and establish a finite-time error bound relative
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית