כתבה
arXiv cs.LG ·
LASER: שיטה חדשה ללמידת חיזוק לא מקוונת
LASER: Latent Space Adjoint Matching for Support-Constrained Entropy-Regularized Offline RL
LASER היא שיטה חדשה ללמידת חיזוק לא מקוונת. היא משתמשת בתיאום משועבח במרחב קודם כדי להשיג למידת חיזוק מרחבית עם הסדרה של זרימה. LASER מראה ביצועים טובים ב-40 משימות OGBench.
תקציר מקורי באנגליתarXiv:2610.08989v1 Announce Type: new Abstract: While offline reinforcement learning (RL) enables policy optimization from static datasets without costly online interaction, it remains bottlenecked by the risk of executing out-of-distribution (OOD) actions. Recent approaches mitigate this by learning a behavior-cloning policy through flow matching and then performing RL within its constrained latent space. However, naively optimizing the latent policy can easily cause the policy to collapse into a brittle mode or exploit sharp artifacts of the learned critic. In this work, we find that entropy regularization is essential in latent-space RL for addressing these challenges. We introduce LASER, a novel offline RL algorithm that applies latent-space adjoint matching to achieve entropy-regulari
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית