כתבה
arXiv cs.LG ·
Linear Recurrent Memory Suffices to Distil a World-Model Policy for Robot Air Hockey
תקציר מקורי באנגליתarXiv:2609.39151v1 Announce Type: cross Abstract: Does memory-dependent control need nonlinear recurrent dynamics? We study simulated air-hockey defence under temporary loss of puck tracking. A DreamerV3 teacher outperforms a memoryless policy under tracking loss, while resetting the teacher's recurrent state sharply reduces performance, which demonstrates that the task requires memory. We distil this teacher into compact recurrent policies with a 64 dimensional state, with a combination of a diagonal linear recurrence and an optional rank-$k$ nonlinear innovation while retaining nonlinear observation encoders and action heads. Across five matched seeds, the purely linear recurrent model ($k=0$) matches both the GRU baseline and the teacher throughout the tested range of tracking loss. Inc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית