כתבה
arXiv cs.LG ·
The World Model Remembers, the Actor Forgets: Dream Rehearsal for Continual Model-Based RL
תקציר מקורי באנגליתarXiv:2607.19749v1 Announce Type: new Abstract: Model-based reinforcement-learning agents of the DreamerV3 family forget catastrophically when trained on task sequences, even when an unbounded replay buffer preserves every earlier experience. We ask a question the continual-RL literature has assumed an answer to but never measured: which component forgets? Under never-clear replay, pre-registered component-level probes (n=3 seeds throughout) show that the world model retains essentially everything measurable about old tasks -- reward discrimination (retention ratio ~1.0), value estimates, and termination structure -- while the actor's behavior collapses. Forgetting in this regime is a channel problem, not a memory problem. We demonstrate this by intervention: with the world model frozen an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית