כתבה
arXiv cs.LG ·
בדיקת תמונה תוך-תפוצה ללמידה עצמית מבוססת-מודל
In-Distribution Imagination for Model-Based Offline Reinforcement Learning
שיפור כושר הדגימה בלמידה עצמית מבוססת-מודל תוך-קו באמצעות בדיקת תמונה תוך-תפוצה. ניתן לראות גם את המודל Gemini.
תקציר מקורי באנגליתarXiv:2609.38673v1 Announce Type: new Abstract: Model-based offline reinforcement learning (MBORL) improves sample efficiency through model-generated trajectories. However, accumulative model error can drive imagined trajectories outside the offline data distribution, leading to unrealistic synthetic data and unstable policy optimization. Many existing methods primarily control rollouts using transition-level uncertainty. We propose \emph{in-distribution imagination} (IDI), a rollout control framework that estimates trajectory support in a learned representation space and adaptively truncates rollouts that leave the offline trajectory manifold. Combined with trajectory-regularized RL, an extension of entropy-regularized RL, IDI consistently improves performance in limited-data settings. Ex
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית