כתבה
arXiv cs.LG ·
Permutation Robustness Is Not Enough: Action Collapse in Multi-Agent Transformer Policies
תקציר מקורי באנגליתarXiv:2610.02848v1 Announce Type: cross Abstract: Transformer policies are attractive for multi-agent robot learning because self-attention can model interactions among agents. However, multi-agent teams are unordered, while transformers typically process agents as ordered token sequences. We study how this mismatch affects cooperative navigation policies under agent-order permutations. Our results show that low permutation error alone can be misleading: policies may appear robust simply because all agents choose the same action. We therefore evaluate policies using both permutation-consistency metrics and action-collapse diagnostics, including action diversity, same-action fraction, and maximum action frequency. A PPO-ID baseline yields non-collapsed behavior but remains order-sensitive,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית