כתבה
arXiv cs.LG ·
Adapting to Changes in Agent Behavior via Finite-Depth Policy Sensitivity
תקציר מקורי באנגליתarXiv:2610.07475v1 Announce Type: new Abstract: Adapting a reinforcement learning policy to changes in another agent's behavior typically requires a large amount of new interaction data. Policy sensitivity provides a first-order prediction of how a locally optimal policy changes with a behavioral parameter, but its computation requires second-order derivatives whose effects propagate across future interactions. We develop a finite-depth framework to estimate this sensitivity by approximating the policy Hessian and mixed derivative using information from a reference environment. The method features an adjustable propagation depth which determines where derivative propagation along the trajectory is truncated. We characterize the derivative contributions omitted by finite-depth propagation a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית