יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת Q להגיעות ב-MEC-Free MDPs

Q-Learning for Reachability in MEC-Free MDPs
אלגוריתם Quasar, הראשון שלא דורש דגימות תנאי, מספק גישה חדשה ללמידת Q להגיעות ב-MDPs חופשי MEC. האלגוריתם, Quasar, נבנה על טכניקת Q-Learning המוכרת, ומשתמש בעדכוני טמפורל-דיפרנס להגיע למדיניות אופטימלית.
תקציר מקורי באנגליתarXiv:2610.01781v1 Announce Type: cross Abstract: Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resultin
קרא במקור המקורי