יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פעולה רפלקסיבית עצמית למשימות קואופרטיביות עם מרחב פעולה רציף

A Self-Evolving Default Action for Cooperative Tasks with Continuous Action Space
מערכת חדשה למשימות קואופרטיביות, המשתמשת בפעולה רפלקסיבית עצמית עם מרחב פעולה רציף. המערכת מסוגלת להתאים את עצמה למשימות שונות ולשפר את התוצאות. המערכת נבחנה במספר תרחישים והוכיחה את עצמה כמערכת יעילה ומוצלחת.
תקציר מקורי באנגליתarXiv:2607.18597v3 Announce Type: replace Abstract: Counterfactual credit assignment has proven effective in multi-agent reinforcement learning (MARL) for discrete action spaces, yet its extension to continuous-action cooperative tasks remains challenging. Existing methods that approximate the counterfactual baseline via Monte Carlo sampling often introduce bias into policy gradients and fail to guarantee convergence to local optima, as the sampled actions may not have been sufficiently trained. To address these limitations, we propose SAFE, a novel MARL framework that employs a counterfactual baseline conditioned on a self-evolving default action sampled from each agent's experience buffer. This design naturally extends to continuous action spaces without relying on additional simulations
קרא במקור המקורי