כתבה
arXiv cs.LG ·
Safety of Latent Communication in Multi-Agent Systems
תקציר מקורי באנגליתarXiv:2609.39788v1 Announce Type: cross Abstract: Latent communication enables multi-agent systems to exchange information directly in internal representation space, reducing the token, computation, and latency overhead of text-based communication. To this end, lightweight trainable links are introduced to map the sender's representations into the receiver's input space. In this work, we show that even benign link training can increase harmful compliance relative to text-based communication while the underlying safety-aligned agents remain unchanged. An attacker can amplify this effect by optimizing the links on harmful query--response pairs or poisoning otherwise benign training data. We further develop a reinforcement-learning attack that rewards harmful compliance alongside benign task
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית