יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת ריפוד עצמאית של סוכנים רב-סוכנים: CASTLE

Independent Multi-Agent Reinforcement Learning with Counterfactual Semantic-Social World Models
למידת ריפוד עצמאית של סוכנים רב-סוכנים, כשהם משתמשים בעולמות-מודל נגד-מציאותיים.
תקציר מקורי באנגליתarXiv:2610.07704v1 Announce Type: cross Abstract: Fully decentralized multi-agent reinforcement learning (MARL), also referred to as independent learning, requires each agent to learn and act using only its local information and experience, without a centralized critic or inter-agent communication. Such a stringent information structure renders the conventional reward signal ambiguous. A poor return may result from an ineffective ego action, an incompatible teammate response, or an effective opponent response, yet scalar rewards alone do not reveal which explanation is responsible. We argue that agents can learn more effectively by prospectively comparing the consequences of candidate actions rather than diagnosing failures only from realized returns. We introduce CASTLE (Counterfactual Ac
קרא במקור המקורי