יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מדיניות זרימה חד-שלבית לתיאום

Flowing Faster to Coordinate: One-Step Online Multi-Agent Flow Policies
נושא המחקר הוא לימוד תיאום התנהגויות מורכבות בין סוכנים רבים. החוקרים מציעים מסגרת חדשה ללימוד פוליסות זרימה יעילות ובעלות יכולת ביטוי גבוהה.
תקציר מקורי באנגליתarXiv:2610.01882v1 Announce Type: new Abstract: Multi-agent reinforcement learning (MARL) provides a powerful framework for learning coordinated behaviors through interactions with the environment. Developing MARL policies requires balancing expressive modeling of complex and multimodal action distributions with efficient training and execution. Generative policies, particularly diffusionbased policies, can faithfully capture complex and multimodal behaviors, but costly iterative sampling hinders their scalability in online multi-agent settings. We propose an Online MARL framework via one-step Flow model (OMAF) that combines expressive generative policies with efficient one-step action generation. OMAF employs a Transformer-based flow policy to capture complex coordination behaviors, while
קרא במקור המקורי