יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מפולת מדיניות סטטית לפריורים תגובתיים מתגברים ב-RLO

From Static Policies to Adaptive Priors in Offline Reinforcement Learning
ה-RLO ה-Offline יכול להשתפר כאשר הוא לומד פריורים מדיניותיים תגובתיים, שיכולים לשפר את עצמם בעתיד.
תקציר מקורי באנגליתarXiv:2609.35880v1 Announce Type: new Abstract: Offline reinforcement learning (RL) has traditionally focused on learning policies for direct deployment under conservative objectives, where uncertainty outside the offline dataset is treated pessimistically to ensure robustness. We argue that this formulation becomes incomplete when an offline-trained policy is subsequently updated through online interaction, as increasingly occurs in modern intelligent systems through test-time adaptation and online fine-tuning. This position paper argues that, in such settings, the objective of offline RL should extend beyond immediate deployment and instead prioritize learning adaptive policy priors: policies that preserve the capacity to improve during subsequent interaction through memory, exploration,
קרא במקור המקורי