כתבה
arXiv cs.LG ·
Space-sampled Value Decay: Forgetting Mechanisms for Non-stationary Reinforcement Learning
תקציר מקורי באנגליתarXiv:2606.11797v2 Announce Type: replace Abstract: Reinforcement Learning agents deployed on physical systems must adapt continually, since degradation and shifting environment conditions change the dynamics (they \emph{drift}) over time. In the hardest version of this problem, the agent interacts with a single system that might drift at every timestep, leaving no opportunity to revisit past conditions -- a setting we call Single Environment, One-Shot Non-Stationary Reinforcement Learning (SEOS-NSRL). We argue that this setting calls for selective forgetting rather than re-learning, and introduce Space-sampled Value Decay (SsVD), which pulls value estimates of randomly chosen elements of the state space to a baseline value, so that outdated information in non visited regions is discarded.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית