כתבה
arXiv cs.AI ·
עבריות והגבלות בבחירת נקודות ציון מאולץ ב-RSSM World Models
Reward Observability and the Limits of Offline Checkpoint Selection in RSSM World Models
במאמר זה, המחברים חקרו את התכונות הסגורות של RSSM World Model, שהוכשר על דגמי אנומליה של LunarLander-v3. הם ניצלו את המודל העולם עבור MPC ו-A2C תוך שימוש באימג'ינציה. הם גם השוו את MPC עם פוליצי BC שהוכשר על הדגמי ההצלחה.
תקציר מקורי באנגליתarXiv:2607.01736v3 Announce Type: replace-cross Abstract: We study the closed-loop properties of a recurrent state-space model (RSSM) world model trained on human demonstrations in Gymnasium's LunarLander-v3. We use the trained world model for zero-shot CEM model-predictive control (MPC) and for actor-critic (A2C) training in imagination. Scored on 100 held-out episodes, the selected model-based A2C policy (trained on world-model checkpoint 280) reaches a mean return of +189.5, matching the best model-free A2C checkpoint (+183.7; 600- and 1000-step episode caps respectively) with ~65x fewer real training transitions. We also compare world-model MPC with a behaviour-cloning (BC) policy trained on the successful demonstrations. The BC policy matches MPC's mean return only under stochastic ac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית