כתבה
arXiv cs.LG ·
חרטה קבועה במשחקים כלליים
Constant regret in general games via higher-order optimism
אלגוריתם חדש למשחקים כלליים מבטיח חרטה קבועה. האלגוריתם, הנקרא HOOD, משלב חיזוי מסדר גבוה עם רגולריזציה אנטרופית. הוא דומה לעבודה עצמאית אחרת שפורסמה לאחרונה.
תקציר מקורי באנגליתarXiv:2609.04113v1 Announce Type: new Abstract: We introduce an uncoupled learning algorithm which, when employed by all players of an arbitrary $N$-player normal form game with up to $K$ actions per player, guarantees $O(N^3\log^2 K)$ individual regret, uniformly over the horizon of play. The proposed algorithm - which we call higher-order optimism with discounting (HOOD) is a variant of optimistic follow-the-regularized-leader (OptFTRL) that combines a discounted $(N+1)$-th order predictor with entropic regularization over a suitable "lifting" of the game's strategy space. This combination of ingredients is purposefully designed to dampen large oscillations of the induced sequence of play in a controlled manner, removing in this way a key stumbling block of previous attempts to achieve c
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית