כתבה
arXiv cs.AI ·
למידה בחלומות, ניצחון באמת: תהליך דינמי רציף למשחק MOBA של עשרה גיבורים
Learning in Dreams, Winning in Reality: A Continuous Dyna Loop for a Ten-Hero MOBA
מאמר זה עוסק בלמידת עולם של משחק MOBA ובשימוש בו להכשרת מדיניות שמנצחת במשחקים אמיתיים. המחברים הציגו תהליך דינמי רציף שמאפשר ללמוד ולהכשיר מדיניות שמנצחת במשחקים אמיתיים. התהליך כולל שלבי למידה והכשרה של המדיניות, והוא משתמש בעולם המודל של המשחק כדי ללמוד ולהכשיר את המדיניות. המחברים גם הציגו תוצאות של ניסויים שהראו שהתהליך הדינמי הרציף מאפשר ללמוד ולהכשיר מדיניות שמנצחת במשחקים אמיתיים.
תקציר מקורי באנגליתarXiv:2610.08033v1 Announce Type: new Abstract: World models are usually judged from the inside: by prediction loss, by the return a policy earns in imagination, or by how convincing their frames look. We judge one from the outside. We learn a structured, multi-agent world model of a complete ten-hero MOBA (206 units, every hero acting every tick, games of up to 6,000 ticks), train a policy only inside it with 1,400-tick free-running imagined episodes, and measure that policy in the real game against the opponent the game ships with. The real game never provides a gradient; it provides the policy's own games as training data for the world model, and an online evaluation that selects and anchors the policy. Run as a continuous asynchronous Dyna loop, the policy wins 70.2% of real games as r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית