יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Turnover-Orthogonal Credit Assignment

Turnover-Orthogonal Credit Assignment for Open-Team Multi-Agent Reinforcement Learning
TOCA הוא שיטה חדשה להקצאת קרדיט בלמידה חיזוקית רב-סוכנים. היא מאפשרת הפרדה בין השפעות פעולות להשפעות החלפת סוכנים. TOCA משתמשת בביקורת מרכזית ובסימנים לאירועים. היא מראה תוצאים טובים יותר משיטות אחרות בסביבות עם החלפת סוכנים תכופה.
תקציר מקורי באנגליתarXiv:2610.02847v1 Announce Type: cross Abstract: Open-team multi-agent reinforcement learning studies cooperative systems in which agents may join, leave, or be replaced during an episode. In such settings, the team return changes both because agents choose useful actions and because the active population itself changes. Standard centralized critics and shared advantages often mix these two effects into one scalar credit signal, allowing surviving agents to be rewarded or penalized for exogenous turnover events outside their control. We introduce turnover-orthogonal credit assignment (TOCA), a value decomposition for open teams that separates action effects, pure turnover effects, and action--turnover interactions. Under exogenous turnover, the event-conditioned value admits a centered de
קרא במקור המקורי