יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת מדיניות אופטימלית ב-MDPs רב-שרשרת: גישה של פירוק היררכי

Average-Reward Reinforcement Learning for Multichain MDPs: A Hierarchical Decomposition Approach
למידת מדיניות אופטימלית ב-MDPs רב-שרשרת: גישה של פירוק היררכי. ניתן להשיג פוליציות אופטימליות ב-MDPs רב-שרשרת עם גישה חדשה של פירוק היררכי. הגישה זו מאפשרת להשיג פוליציות אופטימליות ב-MDPs רב-שרשרת עם גישה חדשה של פירוק היררכי.
תקציר מקורי באנגליתarXiv:2610.10326v1 Announce Type: new Abstract: We study learning optimal policies in average-reward multichain Markov decision processes (MDPs), where the optimal gain may depend on the initial state and recurrence structures vary across policies, creating challenges for reinforcement learning (RL) methods. We propose an asynchronous value-iteration-based RL algorithm that requires no model knowledge beyond the MDP's transition graph and leverages Bather's decomposition to hierarchically partition the state space into communicating subsystems and transient states. This decomposition induces a recasting of the global decision problem into structured subproblems, which our algorithm exploits. We show that the algorithm converges to the optimal gain and produces gain-optimal policies after f
קרא במקור המקורי