כתבה
arXiv cs.LG ·
World-Model Policy Arbiter for Goal-Conditioned Reinforcement Learning
תקציר מקורי באנגליתarXiv:2610.10932v1 Announce Type: new Abstract: Offline goal-conditioned reinforcement learning (GCRL) has produced a diverse set of goal-reaching algorithms, yet no single algorithm performs best across environments, goals, and even different phases of the same task. Rather than deploying only the best-performing policy, we ask whether a set of frozen goal-conditioned policies can be used collectively as a portfolio, deciding at every state which policy should act. Choosing a policy at each state is not straightforward. The policies' own value functions cannot be compared directly: they may use different scales, and some policies have no value function. We need to judge each policy by the states it is likely to reach, even though we can execute only one policy at a time. We also need to a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית