כתבה
arXiv cs.AI ·
SRPO: Setwise Relative Policy Optimization for Multi-Agent LLMs
תקציר מקורי באנגליתarXiv:2609.08452v1 Announce Type: new Abstract: Multi-agent large language models solve complex tasks by coordinating several policies in a shared environment. However, existing reinforcement learning methods usually optimize each response or trajectory separately, even when several outputs jointly cause one state transition. Consequently, the update unit differs from the action executed by the system. To address this problem, we propose SRPO (Setwise Relative Policy Optimization), which treats the active set the minimal set of outputs consumed by one transition, as one multi-agent action. Specifically, SRPO combines member log-ratios into one cardinality-normalized set ratio, assigns one relative advantage, and clips the set once. This formulation unifies division of labor and joint co-ev
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית