כתבה
arXiv cs.AI ·
למידת רפלקסיה באמצעות Q-Learning ללא MECs
Q-Learning for Reachability in MEC-Free MDPs
אלגוריתם חדש ללמידת רפלקסיה בלי לדעת את המצבים הסופיים. האלגוריתם, Quasar, משתמש ב-Q-Learning כדי להגיע למדיניות אופטימלית בלי לדעת את ההסתברויות הקשריות. האלגוריתם נבחן בבסיס הבדיקה הסטנדרטי.
תקציר מקורי באנגליתarXiv:2610.01781v1 Announce Type: new Abstract: Reinforcement learning (RL) for reachability specifications is fundamental to sequential decision-making. Prior work establishes asymptotic convergence to optimal policies, but only through model-based methods that must explicitly estimate the transition probabilities of the underlying Markov Decision Process (MDP). We present Quasar, the first model-free algorithm with asymptotic guarantees for reachability on the fragment of MDPs free of non-terminal maximal end components (MECs), a building block to which every MDP reduces by the standard MEC quotient. Our algorithm follows the classical Q-learning approach, using temporal-difference updates to converge to an optimal policy without ever learning the transition probabilities. The resulting
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית