כתבה
arXiv cs.AI ·
התופעה המופתעת של AVI בשחקנים עצמיים
The Surprising Effectiveness of Approximate Value Iteration in Self-Play
במחקר חדש, AVI (Approximate Value Iteration) הוכיחה עצמה כיעילה יותר מ-AlphaZero בשחקנים עצמיים. AVI למדה פונקציות ערך יותר נכונות, ואף רגישות גבוה יותר לשינויים בשחקן. המחקר חושף את הפוטנציאל של AVI בשחקנים עצמיים, ומעורר תקווה לשיפורים נוספים בטכנולוגיות ה-LM.
תקציר מקורי באנגליתarXiv:2609.09094v1 Announce Type: new Abstract: Combining search with function approximation has driven major advances in game-playing programs, making self-play algorithms more competitive than ever. Still, the computational overhead of the most popular methods, based on Monte Carlo Tree Search (MCTS), can be substantial. In this work, we investigate whether simpler methods remain competitive in non-trivial, moderately sized games such as Connect Four, Hex(7x7) and synthetic games. We train a minimal self-play implementation of Approximate Value Iteration (AVI) and use ground-truth oracles for exact evaluation. Contrary to expectations, our results demonstrate the surprising effectiveness of AVI: it learns more accurate value functions than those learned by AlphaZero, while its one-step-l
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית