כתבה
arXiv cs.LG ·
יסודות תאורטיים ואלגוריתמים יעילים ללמידת סימולטורים
Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning
חוקרים מציעים גישה חדשה ללמידת סימולטורים עבור למידת תגובה מבוססת מודל (MBRL). הגישה מתמקדת בעמידות אסטרטגית במקום דיוק מנבא. המחקר מציג ניתוח תאורטי ואלגוריתמים יעילים.
תקציר מקורי באנגליתarXiv:2605.29032v3 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably exploit minor model inaccuracies, leading to simulator exploitation and a reality gap where policies succeed in simulation but fail in the real world. We propose that the objective for learning simulators should be strategic robustness rather than predictive accuracy, and formulate this as a zero-sum minimax game between a model player and an adversarial policy player. We provide a comprehensive theoretical analysis: (1) an online learning guarantee showing the game is learnable with sublinear regret bounds; (2) a tractable critic-based simplification bounding the global policy-value gap b
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית