כתבה
arXiv cs.LG ·
Revisiting Temporal Regularization for Smooth Control in Deep Reinforcement Learning
תקציר מקורי באנגליתarXiv:2610.07910v1 Announce Type: new Abstract: Deep Reinforcement Learning policies can produce nonsmooth action oscillations that hinder deployment on physical robots. Existing architectural and penalty-based approaches seek spatial smoothness by directly reducing sensitivity to changes in state inputs, but their broad constraints can degrade task performance as stronger smoothing is pursued. Temporal regularization instead constrains action differences along observed transitions, but has been considered unable to provide the spatial smoothness needed under observation noise. We revisit this assumption by proving that the temporal penalty bounds the expected action differences between current states sharing a next state, revealing a spatial effect that empirically extends to spatial smoo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית