כתבה
arXiv cs.LG ·
תכנון תזמון בשלב הבדיקה של קואורדינציה של סוכנים רב-סוכנים על ידי זרימה של גרדיאנט של ערך מחולק
Test-time Multi-agent Coordination by Decomposed Value Gradient Flow
למדנות רב-סוכנים רפורמינטלי בשלב הבדיקה על ידי תיקון פעולה בשלב הבדיקה. המאמר עוסק בפיתוח של תכנון תזמון של קואורדינציה של סוכנים רב-סוכנים, שמשתמש בפונקציית ערך שלמה של סוכנים רב-סוכנים. המאמר גם עוסק בפיתוח של תיקון פעולה בשלב הבדיקה, שמשתמש בפונקציית ערך שלמה של סוכנים רב-סוכנים. המאמר גם עוסק בפיתוח של תכנון תזמון של קואורדינציה של סוכנים רב-סוכנים, שמשתמש בפונקציית ערך שלמה של סוכנים רב-סוכנים.
תקציר מקורי באנגליתarXiv:2610.02554v1 Announce Type: new Abstract: Offline multi-agent reinforcement learning (MARL) faces a persistent trade-off. Expressive generative policies can represent multi-modal coordination in the data, but cannot distinguish high-value regions, while value-optimized policies exploit the learned Q-function but collapse the multi-modal into a single dominant mode. A single agent's mode collapse can break joint coordination, and simultaneous drift across agents can push the joint policy into unseen regions of the action space. We propose scalable coordination via optimal unified transport (SCOUT), the first offline MARL framework to combine a generative foundation model with a learned value function through test-time action refinement. SCOUT trains two decoupled components: a flow-ma
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית