כתבה
arXiv cs.LG ·
Online Change-point Detection for Cooperative Multi-Agent Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.05298v1 Announce Type: cross Abstract: Cooperative multi-agent reinforcement learning (MARL) systems rely on past experience for learning coordinated behaviour, but this experience may become unreliable if the environment or task objective changes during training. In such cases, agents first need a way to recognize that the situation has changed before deciding how to adapt. This paper studies online change-point detection for cooperative MARL using reward-derived signals. We propose \emph{Patterns of Past Rewards} (PPR), a lightweight algorithm-agnostic detector that smooths agents' return streams, highlights recent changes, and applies a statistical drift detector to flag significant shifts. We evaluate PPR in a custom Speaker-Listener environment based on the Multi-Agent Part
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית