כתבה
arXiv cs.AI ·
Q-Shaped Options for Hierarchical Reinforcement Learning
תקציר מקורי באנגליתarXiv:2610.12135v1 Announce Type: new Abstract: Learning to tackle long-horizon, goal-conditioned tasks requires an agent to reason over extended timescales and act across a broad range of states. In principle, Hierarchical Reinforcement Learning (HRL) addresses both challenges through the interaction between action (temporal) and state (spatial) abstraction. First, using an action abstraction to represent temporally extended behaviour as options reduces the effective decision horizon. Second, enabling different state abstractions at each level of the decision process permits greater data aggregation for learning. However, realising these two benefits of a hierarchical policy depends on learning an appropriate action abstraction. Current HRL algorithms fail in one of two ways. Some discard
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית