יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SRPO: אופטימיזציה קבוצתית של פוליצי למערכות סוכנים

SRPO: Setwise Relative Policy Optimization for Multi-Agent Systems
SRPO מאופטימיזה מערכות סוכנים רב-סוכן בעזרת אופטימיזציה של פוליצי ברמת ההחלטה הקבוצתית. המאמר מציג את SRPO כפתרון לבעיות של חידוש פוליצי במערכות סוכנים. SRPO מייצג את הפלטים שמשותפים להפקת תרחיש סטטי כקבוצה פעילה. SRPO מחלק יתרון משותף לכל קבוצה פעילה, ומקליפ את השינוי הפוליצי המשותף. SRPO מסדר את קנה המידה של השינוי הפוליצי המשותף לפי גודל הקבוצה הפעילה.
תקציר מקורי באנגליתarXiv:2609.08452v2 Announce Type: replace Abstract: Multi-agent systems enable complex reasoning and tool use by coordinating agents that divide roles and refine candidate solutions. Existing methods typically update individual agent responses or treat a complete trajectory as one training example. However, these methods may produce misleading policy updates because they assign the same final outcome to responses or trajectory segments that may play different roles in different team decisions. This is because treating each response as an independent update may separate outputs that jointly determine the next action, while treating an entire trajectory as one update may combine decisions made after different observations. These limitations call for a policy update defined at the level of a
קרא במקור המקורי