כתבה
arXiv cs.AI ·
Early Detection of Distributed Backdoors in Multi-Agent LLM Systems: A Characterization Study
תקציר מקורי באנגליתarXiv:2607.24893v2 Announce Type: replace-cross Abstract: Multi-agent LLM systems can be attacked by a payload that no single agent ever holds in full: several poisoned tools each hide one encrypted fragment, spreading them across several agents, and an external step reassembles and executes them after the run. Per-step safety checks that judge each action in isolation may fail to recognize the complete distributed payload. We investigate how early such an attack can be detected while the run is still unfolding, and how robustly it can be caught once its most obvious cues are stripped away. We build a working instance on a hierarchical multi-agent system, run it under benign and attacked conditions across five language models and two tool environments, and record when each fragment is inje
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית