כתבה
arXiv cs.LG ·
עבריות והגבלות בבחירת נקודת-מפגש ב-RSSM
Reward Observability and the Limits of Offline Checkpoint Selection in RSSM World Models
במאמר זה, נחקרים תכונות הסגולגמות של דגם-עולם RSSM, שהוכשר על דגימות אנושיות ב-Gymnasium's LunarLander-v3. נבחן את הפוטנציאל של RSSM לשימוש ב-MPC וב-A2C באימגינציה. נבחן גם את יכולת הדגם לשמש כ-RSSM עבור MPC ו-BC. נחקר גם את ההשפעה של רכיב הפרס של RSSM על עבריות והגבלות בבחירת נקודת-מפגש.
תקציר מקורי באנגליתarXiv:2607.01736v3 Announce Type: replace Abstract: We study the closed-loop properties of a recurrent state-space model (RSSM) world model trained on human demonstrations in Gymnasium's LunarLander-v3. We use the trained world model for zero-shot CEM model-predictive control (MPC) and for actor-critic (A2C) training in imagination. Scored on 100 held-out episodes, the selected model-based A2C policy (trained on world-model checkpoint 280) reaches a mean return of +189.5, matching the best model-free A2C checkpoint (+183.7; 600- and 1000-step episode caps respectively) with ~65x fewer real training transitions. We also compare world-model MPC with a behaviour-cloning (BC) policy trained on the successful demonstrations. The BC policy matches MPC's mean return only under stochastic action s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית