יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

לכיוון אופטימלי של חרפה ב-MDPs עם קונסטריינטים קשים סטוכסטיים

Toward Optimal Regret in Adversarial MDPs with Stochastic Hard Constraints
במאמר זה, המחברים חקרו את האפשרות להשגת חרפה אופטימלית ב-MDPs עם קונסטריינטים קשים סטוכסטיים. הם הציגו אלגוריתם חדש, MA-OPS, שמשלב חיפוש אופטימיסטי למרגין סלטר עם בדיקה פסימיסטית של המדיניות הנבחרת. האלגוריתם מסוגל ללמוד מדיניות עם מרגין גדול של תקינות ולהקטין את החרפה בכל פרק.
תקציר מקורי באנגליתarXiv:2610.12153v1 Announce Type: new Abstract: We study episodic constrained Markov decision processes with adversarial losses under stochastic hard constraints. Specifically, starting from a known strictly feasible policy with margin $d$, we seek to obtain optimal regret while satisfying the expected cost constraints in every episode. In this setting, Stradi et al. (2025) show that a carefully designed mixing rule attains regret of order $\widetilde{\mathcal{O}}(\sqrt{T}/\min\{d,d^2\})$. Interestingly, they also provide a lower bound of order $\Omega(\sqrt{T}/\rho)$ for the same setting, where $\rho$ is the Slater margin of the offline problem and can be much larger than $d$. In this work, we build on their approach to obtain optimal regret dependence on these margins. Specifically, we p
קרא במקור המקורי