יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הכללה מחוץ לתפוצה עם מודלים רצפיים בלמידת חיזוק רב-סוכנים לא מקוונת

Out-of-Distribution Generalisation with Sequence Models in Offline Multi-Agent Reinforcement Learning
חוקרים פיתחו שיטה להכללה מחוץ לתפוצה בלמידת חיזוק רב-סוכנים לא מקוונת. השיטה משתמשת במודלים רצפיים ומאפשרת הכללה טובה יותר למשימות חדשות. התוצאות מראות שיפור משמעותי בביצועים לעומת מודלים קודמים.
תקציר מקורי באנגליתarXiv:2609.03667v1 Announce Type: new Abstract: Generalising to unseen tasks remains a fundamental challenge in offline multi-agent reinforcement learning (MARL). In this work, we present a principled analysis of zero-shot task generalisation in the offline setting and conduct an extensive empirical investigation into the scaling behaviour governing task diversity, dataset size, and network capacity. To facilitate this study, we extend offline sequence modelling architectures to handle multi-task observation and action spaces alongside variable agent counts across tasks. Our primary finding is that scaling task diversity---rather than sheer dataset size is the dominant factor in achieving robust zero-shot transfer. Through large-scale experiments across four challenging environments (Conne
קרא במקור המקורי