כתבה
arXiv cs.AI ·
When Do Options Help? Policy Necrosis and Redundant Coverage in Option-Critic
תקציר מקורי באנגליתarXiv:2609.05508v1 Announce Type: cross Abstract: Option-critic learns options: sub-policies together with a learned rule for when each one hands control back. Its headline result is that performance improves as options are added. We explain that result, with theory and experiment. First, the termination rule option-critic learns by maximising return contributes nothing. When the termination test and the policy that picks options read the same values, the test fires at every step, so the learned rule is identical to always terminating. When that policy explores and the test does not, as in option-critic itself, the rule can block the exploration; there are instances where it suffers $\Omega(T)$ regret while always terminating holds to $O(\log T)$. Forcing termination at every step leaves t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית