יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת ריפוד מרוב-תגמיל

Optimal Multi-Reward Reinforcement Learning
במאמר זה, המחברים חקרו את הבעיה של למידת ריפוד מרוב-תגמיל. הם הציגו אלגוריתם חדש שמאפשר להשיג פוליצי טובה לכל תגמיל, באמצעות שימוש בטכניקות של MVP וגפ-בייסד מולטיפליקטיב ויטס.
תקציר מקורי באנגליתarXiv:2609.36486v1 Announce Type: new Abstract: We study an unknown-transition finite-horizon Markov decision process (MDP) with a finite collection of known reward functions $\{r^1, r^2, \ldots, r^M\}$. The goal is to output an $\epsilon$-optimal policy for every reward using online episodic interaction only. Performance is measured by the policy error $V_{0}^{*, m} - V_{0}^{\widehat\pi^{m}, m}$ where $m\in [M]$ represents the reward function and $V_{0}^{*, m}=\mathbb{E}_{s_1\sim \mu}[V_{1}^{*, m}(s_1)]$. Under this setting, we design a provably efficient algorithm to establish a minimax sample complexity bound of $$ O\left(\frac{SAH^3}{\epsilon^2}\log M \mathrm{polylog}\left(\frac{SAH\log M}{\min\left\{\epsilon, 1\right\}\delta}\right)\right)$$ episodes, with no additional burn-in cost.
קרא במקור המקורי