כתבה
arXiv cs.AI ·
דטקטיב של קולאבורציה בין סוכנים רב-סוכנים
Detecting Multi-Agent Collusion Through Multi-Agent Interpretability
במאמר זה, המחברים פיתחו שיטה לזיהוי קולאבורציה בין סוכנים רב-סוכנים, באמצעות פירוש של תכונות המודל. השיטה, הנקראת NARCBench, נבחנה במספר מודלים, כולל Qwen3-32B, Llama-3.1-70B ו-DeepSeek-R1 32B.
תקציר מקורי באנגליתarXiv:2604.01151v3 Announce Type: replace Abstract: As LLM agents are increasingly deployed in multi-agent systems, they introduce risks of covert coordination that may evade standard forms of human oversight. While linear probes on model activations have shown promise for detecting deception in single-agent settings, collusion is inherently a multi-agent phenomenon, and the use of internal representations for detecting collusion between agents remains unexplored. We introduce NARCBench, a benchmark for evaluating collusion detection under environment distribution shift, and propose five probing techniques that aggregate per-agent deception scores to classify scenarios at the group level, evaluated across four open-weight models (Qwen3-32B, Llama-3.1-70B, DeepSeek-R1 32B, GPT-OSS-20B) and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית