יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אופטימיזציה של תכנון פוליצי עם דגימה פוסטריורית

Offline Policy Optimization with Posterior Sampling
אופטימיזציה של תכנון פוליצי עם דגימה פוסטריורית. פיתוח חדש של תכנון פוליצי שמאפשר חקירה של אזורי תפוצה מחוץ לתפוצה.
תקציר מקורי באנגליתarXiv:2605.07393v2 Announce Type: replace Abstract: A fundamental challenge in model-based offline reinforcement learning (RL) lies in the trade-off between generalization and robustness against exploitation errors in out-of-distribution (OOD) regions. The key to resolving this trade-off lies in enabling the model to explore OOD regions that remain consistent with underlying physical dynamics. However, achieving this is challenging because limited data cannot uniquely identify the dynamics model, and unconstrained exploration is risky. Existing methods often overlook this nuance, addressing the risk through excessive pessimistic regularization, which ensures robustness but sacrifices generalization. To address this, we propose PSPO, which treats the dynamics model as a random variable rath
קרא במקור המקורי