כתבה
arXiv cs.LG ·
Beyond Policy Support: Interaction Constrained Offline Reinforcement Learning for Autonomous Driving
תקציר מקורי באנגליתarXiv:2610.09763v1 Announce Type: new Abstract: Offline reinforcement learning enables reward-driven policy improvement from fixed datasets without requiring online exploration, making it particularly attractive in safety-critical domains. A central challenge, however, is distribution shift: policy optimization may favor actions that are weakly supported by the offline data, rendering value estimates unreliable. Existing approaches primarily control this shift in the policy's own action space. In interactive environments such as autonomous driving, this can be insufficient: a candidate ego trajectory may remain well supported under the marginal behavior distribution while being poorly supported jointly with the surrounding-agent behavior observed in the logged interaction. We refer to this
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית