כתבה
arXiv cs.AI ·
שיפור מודלים עולם חלשים מאחורי סוכנים חזקים
Improving Weak World Models Behind Strong Agents in Atari Pong
חוקרים שיפורים במודלים עולם חלשים מאחורי סוכנים חזקים במשחק Atari Pong. הם בדקו חמישה מודלים: DreamerV3, DIAMOND, TWISTER, Simulus ו-STORM, ומצאו כי המודלים החלשים גורמים לביצועים נמוכים. הם הציעו שיטה חדשה, Concept-Guided Spatial Regularization, לשיפור המודלים.
תקציר מקורי באנגליתarXiv:2607.15142v3 Announce Type: replace Abstract: Strong world-model agents frequently contain weak world models. We study this agent-world-model gap by reproducing five visual world-model agents in Atari Pong: DreamerV3, DIAMOND, TWISTER, Simulus, and STORM, with performance comparable to the reported results, and independently evaluating their frozen world models. First, closed-loop rollout diagnosis qualitatively inspects visual trajectories generated by each frozen model under an independently trained policy. All five models exhibit clear visual or dynamical failures, including ball disappearance, incorrect motion, and invalid ball-paddle interactions. Second, under native zero-shot model-based reinforcement learning (MBRL), a new policy is trained entirely within the frozen model fr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית