כתבה
arXiv cs.AI ·
שבירת תהליכים תחניים בלמידת חיזוק רב-סוכנים
Non-Stationarity Breaks Permutation Surrogates in Multi-Agent Reinforcement Learning: Diagnosis and Remedies
חוקרים בדקו את השפעת התהליכים הלא-תחניים על מודלים של למידת חיזוק רב-סוכנים. הם מצאו שהתהליכים הלא-תחניים יכולים לגרום לתוצאות שגויות. הם הציעו פתרון חדש שמשנה את המודל הריק ומגיע לתוצאות מדויקות יותר.
תקציר מקורי באנגליתarXiv:2604.23716v4 Announce Type: replace Abstract: Reporting guidance for information-theoretic measures is rarely tested against ground truth. We test one guardrail in two multi-agent reinforcement learning games, a social dilemma and a coordination race, where directed influence between selected agent pairs is zero by construction, over 100 seeds. Omitting one precondition, exclusion of the non-stationary training transient, gives false-positive rates of 100.00% and 99.95%: agents annealing exploration independently, in runs that never met, are flagged as influencing one another. Excluding the transient reaches 3.0% in the social dilemma but 11.8% in the coordination game, which stationarity tests explain: 95.7% of social-dilemma series are stationary afterwards against 56.8% of coordin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית